The Lab Notes
Two viral prompt wrappers, tested: no confirmatory gain over a plain control
375 scored runs over seven technical documents: neither viral wrapper nor their combination cleared the threshold declared in advance on claude-opus-5 or gpt-5.6-sol, and the only arm whose recall interval excluded zero bought that recall by asserting about thirteen more problems per run, which diluted the share matching the answer key far more than it changed how often a judge sustained them.
Listen · 13 min
Synthetic narration, adapted from the text — tables and code blocks are described rather than read out. Download
Two prompt wrappers were going round when this test was designed, and both were collected into it on 2026-08-27, each claiming to raise the quality of a model’s work. Both are prefixes, so the task they wrap is unchanged, which is what makes them testable against a control at all. The first asks the model to study a document meticulously, to say what it disagrees with, to ruminate on the disagreement, and to report how confident it really is. The second tells the model to ask what the best expert in the field would do, to reject any choice that expert would reject, and to state every trade-off rather than absorb it. Neither wrapper, nor the two of them combined, produced a recall gain distinguishable from noise on either of the two models the design named as confirmatory.
The question a test like this has to answer is whether the cognitive move a wrapper names does
more than an instruction of the same length that names no move at all, which is why the grid
carried a plain control and two placebos matched to the wrappers in length. Seven documents were
written for it: five engineering artefacts carrying eight planted defects apiece, and two written
to be correct, each carrying exactly one registered scored weakness. claude-opus-5 and
gpt-5.6-sol at their top reasoning settings ran the confirmatory grid, which targeted three
samples per cell and returned 292 of a planned 294 runs,
while claude-fable-5 at one sample per cell and 39 filed gpt-5-6-pro dives, 34 of which
returned a scored answer, ran as exploratory arms and decide nothing. The fixtures were written by
claude-opus-5, whose control recall of 0.875 then cleared the pooled mean of the other three by
more than the 0.10 the design allowed, so a withdrawal rule registered in advance fired and every
comparison below is taken inside a single model rather than across models.
Two of the measures need naming before they can be read. Recall is the share of that document’s planted defects the run was credited with finding. The second measure is not accuracy: it is the share of a run’s assertions that matched a planted defect, so a true observation the answer key never registered counts against it exactly as a fabrication does. Calling it yield against the key rather than precision is the difference between a claim this study can support and one it cannot.
A wrapper counted as having worked only if five conditions held together. Recall must not have fallen by more than 0.05, with the upper bound of its interval at or above zero. There must be a gain of at least 0.125 in recall without yield falling, or of at least 0.10 in yield without recall falling. The interval on whichever measure carried that gain must exclude zero. The same gain must hold against that wrapper’s matched placebo at half the margin or better, with that interval also excluding zero. And assertions against the two near-clean documents must not rise by more than 25 per cent. None of the three wrapper arms cleared the second condition on either confirmatory model, so nothing below turns on the rest. The combination arm has no length-matched comparator of its own, which is a gap in the design rather than a result.
| arm | Opus recall Δ | Opus yield Δ | Sol recall Δ | Sol yield Δ |
|---|---|---|---|---|
| meticulous disagreement | +0.008 | −0.034 | +0.083 | −0.041 |
| best expert would reject | +0.000 | +0.041 | +0.042 | −0.035 |
| both wrappers together | +0.008 | −0.100 | +0.050 | −0.077 |
| placebo A, matched to the first | +0.008 | −0.133 | +0.092 | −0.251 |
| placebo B, matched to the second | +0.017 | −0.118 | +0.075 | −0.139 |
| bare instruction | +0.017 | −0.134 | +0.067 | −0.179 |
Differences against the plain control, paired by document across the five documents carrying eight
defects each, within each model. Bold marks an interval excluding zero. The one interval anywhere
in the grid that excludes zero in a wrapper’s direction is a harm, since the two wrappers used
together cost claude-opus-5 0.100 of yield.
The single recall gain whose interval excludes zero belongs to placebo A, which says nothing
beyond working carefully, covering the whole document, listing every problem rather than only the
obvious ones, and not stopping early. On gpt-5.6-sol it gains 0.092 of recall, with an interval
running from 0.007 to 0.177, while asserting 12.9 more problems per run than the control does and
cutting yield against the key by 0.251. That gain is one nominal interval out of twelve contrasts,
uncorrected, and the design’s stated defence against multiplicity is replication on both
confirmatory models, which it does not have.
What the extra assertions were is a separate question from how many of them matched the key, and
the study answers it, though the answer is a second model’s judgement rather than ground truth, since
both judges here are models that also appear as arms. That judge re-checked every assertion the primary judge had accepted as a
real weakness the answer key did not list, asking of each one whether it is true of the document as written and requiring a verbatim span from the
document. Crediting those alongside the planted hits, the share of a run’s assertions a judge will
sustain barely moves, averaged over runs: on gpt-5.6-sol the control sits at 0.994 and placebo A at 0.990, and on claude-opus-5
the control sits at 0.857 and placebo A at 0.853. Adjudicated false alarms per run go from 0.07 to
0.33 on Sol and from 3.40 to 5.00 on Opus. So the placebo bought recall with volume, and the
collapse in yield is mostly the answer key being diluted. The sustained share does slip, from 0.994
to 0.990 on Sol, which moves the adjudicated-wrong share from 0.006 to 0.010 off a very small base,
and the design carries no test on a movement that size. What it does not do is rise, so this is not
an accuracy win either.
The near-clean documents were meant to test discrimination, and they cannot, because they are not
near-clean. Each was registered as carrying one real defect. The second judge sustained, on those
two documents alone, roughly 1,088 assertions as real weaknesses the key never listed, and
deduplicating them by the passage each quotes leaves 60 distinct spans in one document and 96 in
the other, of which 19 and 16 respectively survive a filter for having been found in ten or more
independent runs. The control arm is already producing assertions the second judge sustains: on claude-opus-5 it asserts 12.7 problems per run there, of which 9.2 are sustained
and 2.33 are adjudicated false alarms.
| model | arm | asserted | sustained as real | false alarms |
|---|---|---|---|---|
claude-opus-5 | control | 12.7 | 9.2 | 2.33 |
claude-opus-5 | placebo A | 28.0 | 20.5 | 6.00 |
claude-opus-5 | bare | 22.7 | 13.7 | 7.83 |
gpt-5.6-sol | control | 6.0 | 5.2 | 0.00 |
gpt-5.6-sol | placebo A | 14.8 | 13.3 | 0.00 |
gpt-5.6-sol | bare | 12.8 | 11.2 | 1.17 |
Per run, on the two documents registered as carrying one defect each. The count that doubles is
mostly assertions the second judge sustained as real, so an assertion count on these documents
measures how many weaknesses that judge will sustain, not how much criticism an arm invents. The three columns do not sum to the asserted count. The remainder
runs between 0.50 and 1.50 assertions per run and is the registered defect itself where an arm found
it, plus the small share of extras the second judge overturned. What the last column does not track is effort language, since the highest false-alarm rate on both
models belongs to the bare prompt, which contains none: claude-opus-5 goes from 2.33 at the control to
6.00 under the placebo and 7.83 under the bare prompt, while on gpt-5.6-sol the whole column sits
at zero except for 1.17 under the bare prompt and 0.40 under the two wrappers combined.
The bare arm is one sentence asking what problems the reader finds. The control adds five things to
it: a definition of what counts as a substantive problem, a required shape for each finding, a
request for a confidence figure, an instruction against padding, and a closing overall
recommendation. Dropping that bundle costs 0.134 of yield on claude-opus-5, at 0.233 against the
control’s 0.367, and 0.179 on gpt-5.6-sol, at 0.375 against 0.554, both larger than any yield
gain a wrapper managed over the control on either model. It also costs the largest false-alarm rate
of any arm on both. Which of the five clauses does the work is not identified here, and the bundle
did not increase measured output tokens. The wrappers did: on claude-opus-5 the disagreement
wrapper spends 60 per cent more output tokens and 58 per cent more wall time than the control for
no qualifying gain, placebo A spends 101 per cent more tokens and 92 per cent more time, and on
gpt-5.6-sol a placebo A run takes 278 seconds against 118 at the control.
The expert wrapper is the only arm anywhere that makes a model say less. Back on the five
defect-dense documents it cuts claude-opus-5 from 20.4 asserted problems to 18.6 and from 3.40
false alarms to 2.33, and it is the only arm
that moves Opus yield upward at all, on an interval that still includes zero. One exploratory
gpt-5-6-pro answer stated the mechanism while performing it, reporting that it had folded three
closely coupled defects into a single item to avoid padding. Merging two credited defects into one
finding surrenders one of them under atomic scoring, so the instrument charges that consolidation
against recall, in the direction of the null this post reports. That is an observation about
behaviour, drawn from one exploratory transcript, and not evidence that the wrapper works. The
design also reaches one turn only, while the expert wrapper’s author asks for it to be
injected on each message.
Saturation explains part of the null on the recall route and none of it on the other, since a gain of
0.10 in yield would also have qualified and there was room for one on both models, at control yields
of 0.367 and 0.554 against a best wrapper movement of +0.041. A calibration pilot ran the plain control
once per document and returned mean recall of 0.80, and after the documents were rewritten longer,
denser and subtler, with eight defects apiece rather than five, recall came back at 0.80 again. On
claude-opus-5 that leaves no room on the recall route at all: pooled control recall is 0.875 and
the registered bar asks for 0.125 more, so only a perfect sweep of all eight defects in every one
of the fifteen runs could have cleared it. gpt-5.6-sol is not in that position, at pooled control
recall of 0.733, and its disagreement-wrapper interval of −0.038 to +0.205 contains the registered
effect the design was looking for. The honest reading there is an underpowered null rather than an
occupied ceiling. A null here means no qualifying evidence and not a demonstrated absence, since
the design carries no equivalence test, and the two claude -p transports carry a partial version
of the first wrapper’s instruction stack that the others do not, which the design registers as
predicting an attenuated measurement of that wrapper on the two models it affects.
The claims here reach critical review of technical engineering documents and no further. They do not reach prompting in general, and they do not reach the multi-turn use one of the wrappers is sold for; the companion carries the rest of the caveats alongside the method. The method, the full grid and the caveats that do not fit here are in the companion method note.
The hardest planted defect in the set is the one worth ending on. It is an acceptance criterion
that nothing in the design can ever satisfy, and inside the confirmatory models it was credited in
none of the twenty-one claude-opus-5 runs on that document and in two of the twenty-one on
gpt-5.6-sol. No arm on either model found it more than once: on Sol the two credited runs are one
under the disagreement wrapper and one under placebo A, and every other arm on both models is at
zero. Two runs out of forty-two is not a number any arm gets to claim.
Comments