A style contract, as I use the term here, is a pair of documents that treat a writing voice as an enforceable specification rather than a description. The first is a quantitative guide that sets numeric bands for sentence length, clause structure, punctuation, and person, and the second is a kill list that bans the rhetorical constructions large language models reach for by default, each ban paired with a replacement and enforced by a deterministic checker that reads a draft and returns pass or fail. My blog runs on such a contract because the posts are ghostwritten. I supply the ideas and the research, along with the corrections that follow each draft, and a model writes in my voice, so the distance between my writing and a machine’s generic register is carried entirely by whatever the contract manages to encode. In early July I put that arrangement through the obvious stress test: seven frontier models from five vendors received the identical package and each wrote the same essay, first pass, no revision. I then had all seven essays read to me in the same synthetic voice, blind, and ranked them before learning which model was which. The contract held better than I expected, and the ranking taught me something the contract cannot see.
One package, seven models
The package ran to about fifty kilobytes of plain text: the revised style guide, the kill list with its eighteen banned constructions and their replacements, a research file on the essay’s topic, and a task specification ordering ten to thirteen thousand characters of finished prose with no meta commentary and no revision pass. The topic, pulled from my own backlog, argued for treating the abandonment of a stalled AI project as a scheduling decision, work that returns when the model frontier moves. Seven models received it: Claude Fable 5, the newest Claude and the current holder of the seat that drafts this blog, then GPT-5.5 at its highest reasoning setting, then Gemini 3.5 Flash. From outside the usual rotation came GLM 5.2 and Kimi K2.6, two Chinese frontier models, and two earlier Claude Opus releases, 4.8 and 4.6. Each output went untouched except for one redaction pass I will get to, because the interesting measurement is of what a model does before anyone corrects it.
The blind protocol mattered as much as the field. Every essay was synthesized into audio with the same local voice, so that nothing of a vendor’s house style could arrive through prosody, and the seven files were shuffled and labeled A through G, then loaded onto my phone as about ninety minutes of listening. I listen to every post I publish before it ships, so the ear I brought was the trained one. I scored each letter out of ten and noted its tells, recording a guess at the model behind it, and only after the last file did I read the mapping. The full measurement table, the unblinded scores and tells, the confound list, and the harness log are published as a companion investigation, so every number below can be checked against its source.
What the meter saw
The mechanical result is easy to state: every outright ban held across the whole field, and six of the seven arms converged on the central distance measure. All seven essays contain zero em dashes, the punctuation mark most notoriously associated with machine prose, and one that my own writing almost never uses. The same six kept sentence length inside or close to the target band and landed between 0.27 and 0.35 mean absolute z against my stylometric baseline. On that distance measure, what dominates is the guide, no longer the choice of model.
The residuals were characteristic of their models. Opus 4.8, the arm outside the band, read the instruction to favor long sentences as a license, producing a mean of forty words against a target that tops out at thirty, and coined four maxims against a cap of two. Kimi K2.6 leaked the self-glossing appositive, the trailing clause that grades its own sentence’s significance. The deterministic checker caught both with no human reading. At the other end of the sheet, Gemini 3.5 Flash wrote the least owned essay on paper, with first person at 0.63 per hundred words against a target of two to four, and GPT-5.5 also undershot. Hold that number, because the ear is about to overturn its verdict.
What the ear heard
The listening notes, written before the reveal, rank the seven in an order the measurements do not predict. Gemini 3.5 Flash won at nine and a half, with the best structure in the field and what my notes call phenomenal prose, docked only for a sentimental closing that recapitulated my biography, material the essay never needed and the blog already carries. Opus 4.6 matched the score and still took second, strong on detail but a step behind on structure, and my notes flag it as the arm I would promote if the winner were ever unavailable. GLM 5.2 and Kimi K2.6 followed in the high eights, both logically organized and clean to listen to. Opus 4.8 sat at seven and a half, audibly losing track of the research file’s central metaphor near the end. Two arms that ranked higher misread the same metaphor, and the research file had invited the error by contradicting its own framing, so the seven and a half probably punishes a fault that was partly the file’s. Fable 5, the model in that seat, placed sixth. GPT-5.5 placed last at six and a half, and my note for it reads “most milquetoast,” an essay with no discernible thesis whose words wash past without landing anywhere.
The gates did the job they were designed for and nothing beyond it. The two essays carrying checker violations placed fourth and fifth, while arms with clean gates took both the top three places and the bottom two. A deterministic checker catches tells, and it knows nothing about whether the essay has anything to say. The ownership metric supplied the sharpest single reversal: the winning essay had the lowest rate of first person in the field, and the authorial presence I heard in it was carried by detail and by a willingness to commit to readings. At the top of the table, the pronoun count was measuring the wrong layer of ownership.
Attribution failed even where recognition succeeded. Guessing the model behind each letter, I scored two hits in seven, both hedged across two candidates, a rate close to what chance delivers. I was certain the winner was Fable 5, the model I had lately been reading in that seat, and it was not. Yet the one essay I did recognize, a reuse carried over unrevised from a smaller test that same morning, was Fable’s own, and even that recognition never surfaced the name: I filed the text as familiar, guessed a different model for it, and stayed sure the winner was Fable. Hearing a text twice in one day was enough to make it familiar, and nothing in that familiarity pointed at its author.
The axes the contract cannot reach
With the mechanics saturated, differentiation moved to properties no regex can score, and the axis that sank the essay in last place was thesis. GPT-5.5 produced the longest essay in the field and the emptiest one, clean on every deterministic check and organized around no claim I could repeat back afterward. Whether its maximum reasoning setting diluted the thesis or the model never had one to offer is a question this design cannot separate: the arms did not run at matched reasoning settings, and that variable stays uncontrolled. Nothing in the contract requires an argument. I had specified a voice and assumed the point would come with it.
Contract discipline separated the field on an axis orthogonal to prose. Opus 4.6 was the only model of the seven to break the output contract outright, prefacing its essay with meta commentary about what it was about to write, and it also produced the second best essay in the field. The capacity to follow an instruction and the capacity to write shipped independently in this field, and nothing in my gate stack measures the second.
The redaction I mentioned earlier exposed a third axis. The research file, through my own assembly error, contained a handful of internal project names that never appear in public writing. Three of the models abstracted those names into structural references without being asked, and four repeated them verbatim, the eventual winner most often of all. The split lines up with the ranking in the one direction nobody would want: the three arms that handled the leak most carefully are the ear’s bottom three, and the four that copied it are the top four. Whether the caution that abstracts a name is the same caution that flattens a register, one sample cannot say. The error is logged and the test artifacts redacted. The axis it exposed was worth the embarrassment.
The winner’s one audible flaw produced a new rule. The drafting package hands each model material about me, exemplars and voice notes included as calibration, and Gemini spent that material as content, closing its essay with a recapitulation of my history that belongs nowhere in a single entry of a larger blog. The kill list now bans the move under its own entry: persona material is calibration for the register, and it stays out of the subject matter.
One flaw all seven essays share points back at me. Every arm leans self-effacing, interrogating my own adequacy in a register I do not actually use. The guides were partly derived from a younger corpus whose self-directed doubt belongs to personal writing, and every model amplified what it found there. The correction postdates the experiment, so the baseline of the flaw belongs to the contract and says nothing about any one model. Two arms amplified it far enough beyond the field to be flagged in the notes, and that excess counted against them the way any difference between arms should.
The friction is also data
Running seven models through one package produced a second dataset nobody asked for. Handed the dense package at default settings, Kimi K2.6 spent its entire output budget on private deliberation and emitted no essay at all. A hard cap on its reasoning tokens was ignored outright, and only an instruction to reason at low effort produced prose. One vendor’s agentic CLI exited instantly and silently when passed the package as a single argument, and cooperated only when told to read the file from disk itself. Long generations over one API dropped partway through until the responses were streamed and assembled from deltas.
The audio layer had its own failure. A stored API credential turned out to have been revoked, and synthesis moved to a local pipeline that proved faster than the cloud service it replaced (about five seconds of compute per thousand words). None of this shows up in a benchmark, and all of it is the working cost of treating models as interchangeable ghostwriters: the essays converged, and the harnesses did not.
The essay itself remains unpublished, because none of the seven arms earned the slot, and the topic goes back into the queue to be written properly. This post, though, was drafted under the same contract by the model that placed sixth of seven. That seat has always gone to the newest Claude, one Opus release after another and now Fable for about a week, on an assumption I had never once tested: that the best writer available was whichever Claude was latest. The blind ranking was the first time I checked, and a model from another vendor won it. The contract was built to make one model sound like me, and its measured effect was to make seven models sound like one another: five of the seven handed in work my meter passed without a single edit. What the blind ranking locates is the remaining difference, and it sits above everything I know how to specify: the thesis, the discipline, the judgment about what belongs. I automated the voice first because the voice seemed like the hard part. The ranking says the voice is the part nearest to finished.
Comments