A worked example · one person and an AI on one record of decisions · 15 September 2026

An Essay on a Throughline

A consultancy's blog essay had been through four drafts in Word. In one evening a person and an AI took it apart into what it claimed and how it said it, checked every claim against a source, and put it back together as a fifth draft. Three facts the four human drafts had carried turned out to be wrong.

Two pagesThis is the account for a general reader. The technical page shows how throughline was used: the five graphs, the style rules borrowed, the renderer, and what the tools taught.
4drafts the author had written before the evening began
15sources opened, read and quoted, one for every dated claim
3factual errors in the fourth draft, found without anyone reviewing it
1evening, of which about ninety minutes was the writing
1of the Word file's 27 parts moves when one character of a paragraph changes; none move when all 108 signatures are deleted

The essay, and the brief

The input was a blog essay for a consultancy's website, about 900 words, arguing that AI differs from earlier technologies for one specific reason. Its author had rewritten it four times in Word. The brief to the AI was two sentences: work out what the essay is arguing and what it is for, challenge and refine that, and then produce an improved fifth draft that is made from the result rather than written afresh. The author added one condition when the plan was put to him: the finished paragraphs themselves had to be part of the record, with their claims, sources and evidence attached, not a document produced beside it.

The first example on this site took a specification to a running service. This one asks whether the same way of working holds when the thing being made is prose.

How the work was divided

The arrangement from the first example held unchanged. The AI proposes and nothing it writes counts until a named person signs it off. Every decision says why it exists. Every claim about the world points to a source that was read, quoted and dated. And the essay a reader sees is produced from the record, so a change to the record changes the essay and nothing else does.

Working out what it was

Nothing was written down until a first pass had recorded why the record would be shaped as it was: what was read, what the subject is, what the work needs, and one decision for each part of the structure, with the alternatives rejected and the reasons. That pass ran twice, because the work turned out to be two things.

The first is the argument: what the essay claims, independent of any wording. Reading the four drafts as a sequence, rather than only the last, was the most useful thing the pass did. What the author had cut, sharpened or replaced between drafts said more about intent than any single draft could. Out of that came a thesis; a set of claims, each labelled as a matter of fact, of definition, of inference or of judgement; the evidence behind each factual claim; the assumptions the argument rests on; the objections a sceptical reader would raise; and a list of things the essay was deliberately not arguing.

The second is the wording: the paragraphs themselves. A paragraph depends on the voice it is written in, the reader it is pitched at and the channel it goes out on. Change any of those and the paragraph changes while the claim does not. So the paragraphs were kept one layer down from the claims, where they can be rewritten for a blog, a social post or a title without the argument moving. That separation is the decision the whole example rests on.

Every factual claim needed a source

The rule was simple: a claim of fact could not be recorded without a link to a source that had been read, quoted and dated. That rule did the work a reviewer would normally do, and it did it before any paragraph was written. Fifteen sources were opened and quoted: a classical dialogue on writing and memory, a space agency's history of its human computers, the founding date of credit scoring, the release date of the first spreadsheet, a cognitive-science paper on search engines and memory, three machine-learning papers, an annual index of model performance, two aviation rules, and two health-service pages.

Three of the fourth draft's facts did not survive the reading. The spreadsheet had arrived "in the 1980s"; the source says October 1979. The first language models had turned out to "write working code"; the paper says the base model solved none of its benchmark and managed only simple programs. A laboratory "produces the result" and a clinician "decides what it means"; the health-service page supports less than that, so the claim was scoped down to what it says. Nobody caught these by reviewing. Each was caught because the claim could not be recorded without a source, and the source could not be linked without being read.

Two more things were recorded rather than fixed. "Each was feared at the time" is evidenced for one of the five technologies, so it became an inference resting on a stated assumption. And the essay's strongest sentence, that the model "does not think, assess or choose", rests on a definition of thinking that is now an assumption of its own, which the essay can choose whether to state.

The sceptic gets a seat

Good persuasive writing acknowledges the strongest objection openly and answers it. The fourth draft did not. So the objections were made part of the record, with a rule that an objection nothing answers is an error. Four were written in the sceptical reader's voice: that current models do reason; that keeping a person in the loop is a passing phase; that this is the usual moral panic; that earlier tools were opaque too. Each is now answered by a named claim, and the fifth draft gained one new paragraph and two new sentences that answer them in the copy.

Putting it back together

Twenty-three paragraphs and six smaller pieces, the title, the subtitle, the teaser for social media, the sign-off, were written, each connected to the claims it carries, the evidence it names and the house-style rules it follows. A paragraph that carried no claim was an error; so was a title or a teaser that carried none.

That is how a drift in the fourth draft came to light. Its social-media teaser said the essay compared AI "with the similar things", while the essay's own opening argues that the usual comparison is the wrong one. The second and third drafts had it right; the fourth had broken it. The teaser could not be connected to the claim as written, so it was restored.

The house rules did their own catching. The sign-off was in italics, against a convention that avoids them. The subtitle read "a assessment". An abbreviation was never spelled out on first use. Each was a rule the text could not honestly claim to follow until it changed.

The same essay, every time

The result is not a fifth draft in the usual sense. It is a small program that reads the record and writes the Word file, in the house format, with every timestamp inside the file fixed. The fifth draft is the record, laid out for a reader. Change a claim, and the paragraph that carries it is flagged until it is redone; change a house rule, and every paragraph that follows it is flagged too.

What that is worth needs stating carefully, because "the same input gives the same output" is true of almost any program. Two things were measured on the finished document. Change one character of one paragraph's prose, and the document changes in exactly two places: the copy a reader gets, and the one part of the Word file that holds the prose. The other twenty-six parts of that file do not move. That was tried on all twenty-eight paragraphs, one at a time, and all twenty-eight behave the same way, so no sentence in the document is unaccounted for. Then remove every signature from the two records the document is built from, 108 of them written as 216 lines, and nothing moves at all.

This page used to carry a different figure: zero bytes changed when the essay was regenerated after everything had been signed off. The author's response to it was that it is not very impressive, and he was right. It proved nothing. The program reads an item's words and never reads which person accepted it, so the figure would have read zero whoever signed, and whether or not anyone signed at all. It was going to be zero before anyone ran it. Sign-off is held in the record instead, where a check names any item whose words have moved since a person accepted them, as a warning in the ordinary check and an error in the strict one. The two measurements above replaced it on 17 September. Both are written down as assertions in a test, which goes red if either stops being true, though nothing yet runs that test on a schedule. The scope is worth being exact about, because a wider claim would be the same mistake again: it is a paragraph's prose that moves the document. A paragraph's reasoning, its title, a section's description and a claim in the argument behind the essay all leave the document exactly as it was, which is the separation the whole example rests on.

Two things the essay knowingly does against the rules were recorded as such rather than hidden. The house style asks for fifteen to twenty words a sentence; the essay averages twelve, because its argument is carried by short declarative sentences. And the budget said 900 words; the fourth draft's body was 1,095 and the fifth is 1,169. Both were written down as deliberate deviations with their reasons, for the author to accept or overturn.

The first page of the generated Word document: title, teaser, byline and the opening paragraphs.
The first page of the generated document. The essay itself is the company's unpublished copy, so this is a thumbnail rather than the text.

Where the evening went

Most of the evening was not the writing. Signing off 157 decisions exposed three faults in the tools, and each was fixed and released before the evening ended. The sign-off screen failed on every record that borrowed from another; the fix was one line, and a test that should have existed. Connecting 153 references to the borrowed sources took about fifteen minutes, because each one re-checked six sources over the network; that was measured rather than guessed, and reported with the numbers. And the command handed to the author for signing off two items at once did not work, because the command took one at a time. He asked whether it should take several. It now does.

Then the author declined to press the button 157 times. He delegated the signatures in writing, and the AI signed on his instruction with the delegation recorded in every change. Whether that is a sign-off or a formality is the honest question this example leaves open.

Who did what

Only the person couldOnly the AI wouldBoth, through the record
Write four drafts, and insist the finished paragraphs be part of the recordOpen fifteen sources in one sitting and quote each with its dateTurn "the 1980s" into 1979 without anyone arguing
Choose which house style the essay followsRead the drafts as a sequence and treat the changes as intentKeep the argument free of any wording, so the wording can change
Catch a hand-over command that did not runMeasure a slow tool instead of guessing, and report the numbersRecord a knowing deviation as a note the author can overturn
Decide that 157 confirmations were performative, and say soFix and release three tool faults the same evening they were foundRetire a figure that proved nothing, and put a check in its place

What went wrong, and what caught it

  • Section notes and a maintainer's comment leaked into the first output. Caught by reading it. Fixed, along with the tests that should have come first.
  • A hand-over command that passed two items to a tool that took one. Caught by the author running it. Fixed in the tool the same evening.
  • The sign-off screen was broken for any record that borrowed from another. Caught by the author on first real use. Fixed and released, with the check that should have caught it.
  • Two commands run at once corrupted each other's reading. Caught when one came back empty. Fixed by running them one at a time.
  • A release made on the wrong repository, because a command named none and two shells raced. Caught by reading the address the command printed; deleted before anything published. Every such command now names its repository.
  • A figure on this page proved nothing. Zero bytes changed when the essay was regenerated after sign-off, which was never capable of any other value, because the program does not read who signed. Caught by the author reading his own example page two days later. Replaced by two measurements that can fail, and by a rule that a published figure must be one that could have come out otherwise. Measuring it properly found five faults in the program, including a paragraph marked rejected that was still being published with every check green.
  • The evening ran long. The author said so. Most of it was the toolchain, not the writing; the writing was about ninety minutes.

Time

PhaseThe AI, elapsedThe person, attention
Working out the argument, reading the sources, writing the paragraphs, first generated draftabout 1 h 30two decisions, by chat
Challenging the result and connecting every referenceabout 0 h 45none
Tool faults: diagnosis, fixes, three releasesabout 1 h 30two sign-offs, three deployment approvals
Signing off the record, assessmentabout 0 h 20one instruction

The AI's hours are from timestamps and the session record, and include waiting on its tools. The person's attention was nineteen messages, most of them under a line.

What this shows

Wording is presentation. Keeping the paragraphs one layer below the claims is what let one argument carry a blog, a teaser and a title without drifting.
Make evidence a precondition, not a courtesy. One rule found three errors that four human drafts had carried. Nobody reviewed anything; the claims simply could not be recorded until the sources were read.
Read the drafts as a sequence. What the author cut and sharpened between versions showed which claims were load-bearing better than any draft alone.
A record is not a document. The tool's own output is right for a reviewer and wrong for a reader. The gap was closed with a small, tested program, not by hand-editing the output.
A deliberate deviation is worth recording. A rule the text knowingly breaks is written down with its reason, never claimed as met; the author decides.
Measure the slow thing. Three seconds a step, 1.3 of them network, was worth reporting; "it feels slow" was not.
Sign-off has a price, and the person sets it. When the signatures became performative, the author said so and delegated them. The record shows the delegation; what it cannot show is whether the reading happened.