Vibe researching, as I use the term here, is directing frontier models to research a soft problem: one where the number you want is published nowhere, and an answer has to be assembled by inference from scattered public statements, figures that are reported but never audited, and a handful of disclosed anchors. The coinage follows vibe coding, and it inherits the analogy’s whole shape. In vibe coding you describe what you want, accept what the model produces, and keep moving as long as the thing appears to run, and the documented failure mode is shipping code nobody read. Vibe research works the same way, with one difference that governs everything downstream: research has no compiler. Wrong code eventually crashes into a runtime that does not negotiate, but a wrong number in a fluent report just sits there, formatted, cited, and confident, waiting for a reader to believe it.
The name, it turns out, is already partly occupied. By mid 2026 “vibe researching” has a life in the literature, where it names AI agents running the mechanics of a conventional research workflow (literature review, experiment design, analysis, drafting) while a human supervises scope and judgment, a usage with its own arXiv preprints and a workshop track. I am keeping the word for a harder case, and the difference matters: the published sense automates research whose answers can in principle be checked against the world, while the problem here has no checkable answer at all, so the inference is not a labor saving but the entire product. This blog has circled the word once already, when post 047 audited five years of my trading instincts; that post examined the vibes, and this one keeps the name for the method.
The target was the serving margin on frontier inference: of every dollar a lab charges for tokens at list price, how much survives the marginal cost of serving them. No lab discloses this. What exists in public is a scope-mixed scatter of analyst claims (Dylan Patel’s “north of 80%” serving-margin floor, a SemiAnalysis 75 percent API gross-margin model), a 90 to 95 percent figure about an unnamed lab, a few reported numbers nobody outside can audit, and exactly one clean disclosure anchor, the throughput and cost data DeepSeek published in early 2025. The result of the week this post describes is live at margins.ashitaorbis.com: an interactive calculator, a long report, an evidence board, an MCP connector that gives agents read access to the same engine, and a feedback intake. Its central scenario lands at roughly 77 percent, and that number is a derived analyst synthesis, labeled that way on the page itself. Nothing in this post claims a measured fact about any lab’s books, and a unit serving margin is not a company gross margin under any reading.
The thesis is the one the project forced on me, and it extends a pattern this blog has already paid for twice. The first working version took about a day. Making the research underneath it survive review took the following four days, six methodology releases, three gate verdicts of NO-SHIP, and a defect ledger that closed at 45 findings fixed (none rejected). By the end the verification apparatus, and never the calculator, was most of the artifact. Post 035 measured a revision tax of roughly two to one on a written investigation, and post 039 found that the hard part of checking facts is deciding which findings are real rather than generating them. Vibe research runs the same inversion one level up, at a worse ratio, because the thing under review is no longer prose about an artifact but the epistemic status of every number the artifact contains.
The prototype day
Version 1 went live on July 9, and the day itself is the honest advertisement for the method. Grok swept X for margin claims, because it is the one tool in the roster that can see X at all. GPT-5.6 Pro ran the provider deep dives, half an hour of sustained reasoning about each lab’s economics. Claude subagents built the cost engine and drafted the report, and an orchestrator stitched the output into an analysis in ten sections with the calculator on top. Analyst statements, a hardware table, provider case studies, a research annex: a volume of assembly I would once have budgeted weeks for arrived, as with the investigation in post 035, before dinner.
The prototype also carried a credential that made it feel finished. DeepSeek’s disclosure is the one place a frontier lab has published enough operational detail to reconstruct a serving margin (84.5 percent at list prices), and my calculator’s replay of that disclosure produced 84.6. A model of an undisclosed economy that reproduces the field’s clearest disclosure-derived anchor to a tenth of a point reads as validation. I read it that way, and I was wrong on both digits.
The match was two errors
The agreement was an accident. The disclosed throughput the replay leaned on, 73,700 input tokens per second per node, includes DeepSeek’s 56.3 percent disk cache hit share, so treating it as a fresh prefill benchmark inflated the modeled efficiency. A redundant utilization divisor elsewhere in the replay deflated the result by almost exactly the same amount. Honest accounting, with the cache share stripped out and the divisor removed, reproduces the disclosure-derived benchmark at 84.0 percent. The two errors were invisible precisely because they canceled, since every auditing instinct keys on numbers that look wrong, and these conspired to look perfect. If the project has one exhibit for why vibe research needs adversarial verification rather than plausibility checks, it is a broken model matching the truth to a tenth of a point.
Review began within a day of publication and returned the verdict that set the week’s course: two separately run model reviews (a GPT council of four personas and a separate GPT-5.6 Pro adversarial pass) agreed the thesis was sound and the presentation was, in their words, “publication-blocking”. The same round surfaced the defect class that vibe coding produces when nobody reads the diff. Switching analyst perspectives in the calculator silently overwrote each provider’s cache read tariff with Anthropic’s 10 percent (the selected models billed cache reads at Grok 25 percent, Kimi 20, GLM 19, DeepSeek under 1). The sensitivity range meant to guard against lens shopping had unioned incompatible cost lenses and printed an endpoint of minus 867 percent for one model. Neither number was hard to fix; what was expensive was everything required to find them.
Four models across three vendors, one honesty problem
Directing, in the vibe coding analogy, suggests one person prompting one model, and the week’s working reality was closer to a small newsroom of four models drawn from three vendors, two of them Anthropic’s own. The intended roster was two, not four: Fable, Claude’s frontier tier, for the research and the building, and GPT-5.6 Sol for the reviews. The other two seats were conscriptions: Grok holds one because no other model in reach can see X, where the analyst claims live, and Opus holds the other because Fable’s safety classifiers kept tripping on legitimate work. So an Opus orchestrator plans, routes, integrates, and runs the gates, deliberately barred from bulk execution so its context stays free for judging what comes back; Claude subagents do the building and drafting. GPT-5.6 review gates sit at every milestone, with plan reviews before code and final gates before deploys, under a standing rule that every remediated finding is verified by a fresh GPT instance that never saw the fix argued for. Grok owns the X sweeps, and GPT-5.6 runs in two configurations, Sol at the gates and Pro on the provider dives and the outside reviews.
The arrangement began as a workaround. Early in the arc, Claude’s safety classifier flagged the project’s own browser test suite, a benign set of input validation checks (render the page headless, refuse malformed share links), because the suite’s original framing described forged payloads against production, and that framing alone was enough to trip a downgrade mid-session. The permanent fix routes anything whose description pattern matches offensive security to GPT directly, while the suite itself is now described as what it is, input validation run locally against a file build. The anecdote earns its place because it shows what the roster actually is: coverage rather than redundancy, with each vendor contributing the thing it alone can see or will touch, X access from one, long adversarial reviews from another, bulk construction from a third.
The gates earned their cost in verdicts. The council reviewing the first improvement release returned NO-SHIP with a surgical fix list. The next release failed its final gate too, which found, among other defects, a permalink forgery path: a crafted share token declaring one schema while carrying another could route through a legacy migration and render 71.6 percent under a locked replay’s name. Nine simulated reader audits, five checking the site’s faithfulness to the specific public figures it cites and four running anonymous reader archetypes, found sixteen defects severe enough to block publication that two full review rounds had missed. The blind council gating the redesign refused it unanimously before the remediation round cleared. The regression suite grew in proportion, from 29 assertions at first release to 499 engine assertions, 220 contract assertions, and a browser harness that replays the attack payloads the reviews invented.
What a verified vibe looks like
What verification kept producing was machinery more than corrections. The evidence base grew from 18 claims to 34, and the growth was the cheap half. The expensive half was the taxonomy every claim now carries: a scope axis recording whether a number describes a token SKU, an API product line, a paid user cohort, a company gross margin, or a business segment, and a provenance tier running from audited disclosure down through reported figure, quote known only through a clip, and analyst assumption. The registry bins claims, not people. A floor is never rendered as an interval, so Patel’s “north of 80%” shows as compatible with every bucket at or above 80 rather than as sitting inside any one of them, because a floor is not a ceiling. The bucket at or above 90 holds exactly one exploration route, and it is a page-authored owned-fleet procurement reconstruction, not the batch size mechanism quoted above it: the engine’s own reconstruction of that mechanism tops out at 87 to 89 percent at market rates, short of 90. Other procurement reconstructions that also cleared 90 were pruned as redundant, leaving the one. The same discipline overruled me directly: my own sketch had the range explorer driving the calculator, and a design council rejected it because a headline that can be steered from a conclusion reopens conclusion shopping, so the causal arrow runs one way and the evidence board can never set the number.
Provenance work cut in both directions, and the honest record keeps both. One figure the report leans on, the claim that Anthropic’s inference infrastructure margin moved from 38 percent to above 70, had been carried for days as reported but unverified through X relays, under a registry rule holding misattributed figures at second tier because the source newsletter was paywalled. The margin sentence turned out to sit in the newsletter’s free portion. A direct fetch archived the primary text verbatim (curly apostrophe and all), and the record upgraded to primary provenance. The opposite case stands beside it: three of the X posts the routes and claims rest on are unretrievable for a durable capture, because Grok could transcribe them during the dated sweep but no other automated path or archive could reproduce them, so their durable provenance is that sweep transcription rather than a live capture, and the page says so plainly instead of papering it over. Every X quote that was adopted got verified by machine, character for character, against the sweep file. One blockquote that had drifted in transit, “overage” quietly normalized to “average” with two lines dropped, was restored to its original, typo intact and marked sic.
Adjudicate before you remediate
The strongest test came on July 13, when a GPT-5.6 Pro instance reviewed the live site cold, holding no internal documents and no context beyond what any reader sees, and returned eighteen findings across substance, epistemics, and presentation. Post 039’s lesson was that findings are cheap and deciding which findings are errors is the work, and the outside review reproduced that lesson at research scale. Adjudication ran finding by finding against what the site already says. It sorted the eighteen into six valid (every one curable in a sentence), eight already mitigated by caveats welded beside the very numbers the reviewer quoted, two that relitigated deliberate design decisions, and two already tracked in the backlog. One adjudicating agent did not argue at all: it ran the engine to check the review’s objection to the site’s own claim that new hardware guarantees a high margin for any frontier model, and produced a counterexample the site itself contains, a model whose list price never enters the 90 to 95 zone on a GB300 fleet at any utilization, so the universal got bounded and the counterexample now appears in the text.
The arithmetic of that adjudication is the post’s whole argument in one row. A cold expert review of a heavily reviewed artifact was still worth commissioning, because six real defects had survived extensive prior review, including nine simulated-reader audits, two of them plain currency errors (a chip count the site’s own hardware table contradicted, and a launch date off by three weeks). And remediating the review blind would have been destructive, because eight of the eighteen demands were already met by text the reviewer had no way to weight, and two more would have reversed choices the design made on purpose. The valid fraction in this review, one third, is the planning prior a verification process has to budget for, and the reason adjudication is a stage of the pipeline rather than a courtesy extended to the author’s feelings.
The tax, one level up
Post 035 priced the revision tax on a written investigation at roughly two to one, revision over generation, and treated the inversion as the uncomfortable news, since generation had become the cheap part of research. This project prices the same inversion one level up, and the ratio moved. One day of generation bought four days of verification and six releases in which the numeric engine mostly held still, because the reviews verified the arithmetic early and then kept failing the presentation, the provenance, and the identity of the numbers instead. It also bought a codebase in which the calculator is a minority stakeholder beside its own regression suites, typed claim registry, permalink codec, and quote verifier. And the daily count understates the price, because the days were not equal in width: generation ran largely single file, while verification fanned out into councils, simulated readers, and fresh-verifier checks that billed in parallel. The token ledger, reconstructed from the week’s session transcripts, puts the imbalance near twenty to one, in a defensible range of roughly fifteen to twenty-five: the measured Claude-side core alone runs eighteen to one, 527 thousand tokens of generation against 9.3 million of verification, the measured GPT review layer lifts it to twenty-one, and the estimated tails for the Pro dives and councils, the loosest numbers in the ledger since the Pro interface reports no token telemetry, nudge the center higher without changing the story. On the wall clock the tax was one day buying four; in tokens it was one buying twenty. The tax changed what it purchases. On post 035’s investigation it bought corrected prose; here it bought machinery that makes the next claim cheaper to verify than the last, and that machinery, more than the 77, is the asset the week produced. And the tax kept collecting after I thought it had stopped: two days past the last of those six releases, while I was drafting this, a seventh caught an ordinary shared link silently rewriting its own margin from 77.9 to 92.3 percent, fourteen points on a copy-paste that every prior gate had passed.
The last check belongs at the end because everything above is models checking models. The site’s own ship gate, defined before the redesign was built, was a comprehension test: five human readers, ninety seconds each, passing only if four could say what the number at the bottom of the page is without being led into reading it as the company’s profit margin. The formal gate never ran; what ran was its casual form. I shared the site with a first round of real readers, and what came back was not a stack of findings to adjudicate but the one thing the rest of this record never produced: good vibes. No misreadings were volunteered, no defect reports came back, just readers who followed the argument and liked the thing. Friendly readers on their own time are a softer instrument than five strangers on a stopwatch, but they are the audience the site was built for, and their reading points the right way. The generation took a day, the verification took the week, and the final check was the oldest kind, actual people reading the thing. It seems right that a post about vibe research signs off on the vibes of its readers.
Addendum: the update system
A number assembled by inference decays: prices move, disclosures land, hardware generations turn over. Since first publication the site has had a maintenance loop to match. A daily sweep reads X through Grok, the provider pricing pages, and the news surface, and only files what it finds into a review queue; nothing edits the site on a sweep’s word. A weekly release applies queue items the author has approved, through the same gate that shipped every version before it, full regression suites and then deploy, with an expedited same-day path reserved for anything that would invalidate a headline number. The public repository at github.com/AshitaOrbis/inference-margins mirrors each release through an allow-listed publish script rather than exposing the working tree. Among the queue’s first entries was new analyst discussion of DeepSeek’s revenue run rate; it entered as a finding, waited for approval, and shipped through the gate rather than being pasted in, which is the point. The verification machinery this post described is what makes maintenance affordable: each new claim rides the same registries and suites as the last, and a static report about a moving economy is just a wrong report on a delay.
Comments