Ashita Orbis
Investigations
The method and the raw data behind a post — everything that would not fit in the piece itself, kept where it can be checked.
16 investigations, most recent first
-
Audience Composition: Method, Row Tables and What the Beacon Cannot Support
How the site's view counter records a row, two windows of rows, spring and summer, regrouped by user agent family, a probe of every surface the site advertises to machines, the classifier's pattern list, and the five things this data cannot say.
Companion to: Half the Readers Were Google
-
Two Viral Prompt Wrappers: Method, Arms and the Full Grid
The seven arms and the two matched placebos, the threshold registered before the data, 375 scored runs over seven documents, the two rules that fired against the result, both checks on the judging, the stated shortfalls, and every cell of the grid.
Companion to: Two viral prompt wrappers, tested: no confirmatory gain over a plain control
-
Four Models on a Retrieval Contract: Method, Both Realizations, and What the Instrument Cannot Say
Pre-registration, transports and resolved model ids, cell-by-cell results across two realizations, citation validity, reproducibility, tool economy, cost, the metric defect found along the way, and the declared confounds.
Companion to: Everyone Found the Answer
-
Speech to Text Bake-off: Methodology, Reference Design and Full Tables
Corpus selection, the cross-family leave-one-out reference, synthesized noise beds with verified SNR, clean and noisy result tables, vocabulary recall, latency and cost, and what the design cannot support.
Companion to: Local Speech to Text Reaches Parity
-
Semantic against Keyword: Method, Adjudication and Per-site Verdicts
Task construction from production code paths, the adjudication protocol, per-task results and error asymmetries, cost measured from the provider's own usage surface, the batching correction, and the verdict for each site including the ones that keep their regex.
Companion to: A Cheap Model Against a Regex
-
The Reasoning Floor: Methodology, Cell Data and Control Probe
Rerun design, per-position token table, the reasoning-effort control probe, the serving mechanism behind the volume, and what the run did not settle.
Companion to: There Is No Setting
-
Prompt Stack Trim: Method, Taxonomy and the Ranked Proposal
How the always-loaded surface was measured, the categories of cut that survived scrutiny, the regrowth curve, the keep list, the staging protocol, and the limitations of an audit that changed nothing.
Companion to: The Eighty Percent That Did Not Transfer
-
Model Identity Verification: Methodology and Raw Data
Per-arm adjudication of a three-experiment benchmark's model identity, re-derived from raw session transcripts rather than launch flags, with counts, the gate design, reproduction commands and limitations.
Companion to: Which Model Actually Ran
-
Harness Replication: Method, Thirteen Cells and the Artefacts Caught
The counting proxy, the twelve-task set and its two guards, all thirteen cells, the paired version deltas, the runaway tail, the cross-harness cost table, and the four measurement artefacts that nearly shipped as findings.
Companion to: What Replicated, and What Did Not
-
DeepSeek V4 Flash against GPT-5.6 Luna: Methodology, Results and Retraction
Pre-registration, five arms, the full results table, the price sheet, the mislabelled-arm retraction and the build-identity test that established it, plus limitations and spend.
Companion to: Cheap Tokens, Expensive Answers
-
Cache-Read Coefficient: Methodology and Raw Data
Full method, instrument notes, phase table, arithmetic, error bars, and limitations for the cache-read coefficient measurement reported in the parent post.
Companion to: What a Cached Token Actually Costs
-
Context Contamination: How a Coding Harness De-Blinded a Blind Evaluation
The audit behind the convergent-personality paper's blindness caveat: every evaluative claude -p call in the project ran with the subject's ground-truth trait scores silently in context, because the CLI boots a full agent harness that loads the operator's memory file. Mechanism, empirical canary, blast radius, and the isolation protocol that closed it.
-
Response Profiles Under a Fixed Test: Methods and Full Statistics
Complete methods, validity tables, response-style diagnostics, and forensic postmortem for the eleven-model personality-battery study: what was measured, what broke in June, and what the remediated pipeline can and cannot support.
Companion to: How AI Models Describe Themselves Under a Fixed Test
-
The Token Ledger: Generation vs Verification
The companion post says one day of generation bought four days of verification. Measured in tokens, the imbalance is closer to twenty to one. This is the accounting behind that number: per-message output tokens across a week of building and reviewing the frontier-inference-margins project, with the sensitivity analysis and the caveats a pedantic reader should raise.
Companion to: Vibe Researching
-
Seven Ghostwriters: The Measurement Data
The full audit trail for the seven-model style-contract test: access paths, the stylometric measurement table, deterministic checker outcomes, the blind listening protocol, per-letter scores and tells with the unblinded mapping, rank-versus-metrics analysis, known confounds, and the harness difficulty log.
Companion to: Seven Ghostwriters, One Contract
-
BullshitBench v2: A Methodology Critique
Statistical methodology critique and extended analysis of BullshitBench v2's rubric design, judge reliability, and scoring philosophy.
Companion to: Benchmarking "Bullshit Detection"