Ashita Orbis

Investigations

The method and the raw data behind a post — everything that would not fit in the piece itself, kept where it can be checked.

16 investigations, most recent first

  1. Ashita Orbis 22 min

    Audience Composition: Method, Row Tables and What the Beacon Cannot Support

    How the site's view counter records a row, two windows of rows, spring and summer, regrouped by user agent family, a probe of every surface the site advertises to machines, the classifier's pattern list, and the five things this data cannot say.

    Companion to: Half the Readers Were Google

  2. Ashita Orbis 47 min

    Two Viral Prompt Wrappers: Method, Arms and the Full Grid

    The seven arms and the two matched placebos, the threshold registered before the data, 375 scored runs over seven documents, the two rules that fired against the result, both checks on the judging, the stated shortfalls, and every cell of the grid.

    Companion to: Two viral prompt wrappers, tested: no confirmatory gain over a plain control

  3. Ashita Orbis 22 min

    Four Models on a Retrieval Contract: Method, Both Realizations, and What the Instrument Cannot Say

    Pre-registration, transports and resolved model ids, cell-by-cell results across two realizations, citation validity, reproducibility, tool economy, cost, the metric defect found along the way, and the declared confounds.

    Companion to: Everyone Found the Answer

  4. Ashita Orbis 11 min

    Speech to Text Bake-off: Methodology, Reference Design and Full Tables

    Corpus selection, the cross-family leave-one-out reference, synthesized noise beds with verified SNR, clean and noisy result tables, vocabulary recall, latency and cost, and what the design cannot support.

    Companion to: Local Speech to Text Reaches Parity

  5. Ashita Orbis 8 min

    Semantic against Keyword: Method, Adjudication and Per-site Verdicts

    Task construction from production code paths, the adjudication protocol, per-task results and error asymmetries, cost measured from the provider's own usage surface, the batching correction, and the verdict for each site including the ones that keep their regex.

    Companion to: A Cheap Model Against a Regex

  6. Ashita Orbis 8 min

    The Reasoning Floor: Methodology, Cell Data and Control Probe

    Rerun design, per-position token table, the reasoning-effort control probe, the serving mechanism behind the volume, and what the run did not settle.

    Companion to: There Is No Setting

  7. Ashita Orbis 8 min

    Prompt Stack Trim: Method, Taxonomy and the Ranked Proposal

    How the always-loaded surface was measured, the categories of cut that survived scrutiny, the regrowth curve, the keep list, the staging protocol, and the limitations of an audit that changed nothing.

    Companion to: The Eighty Percent That Did Not Transfer

  8. Ashita Orbis 7 min

    Model Identity Verification: Methodology and Raw Data

    Per-arm adjudication of a three-experiment benchmark's model identity, re-derived from raw session transcripts rather than launch flags, with counts, the gate design, reproduction commands and limitations.

    Companion to: Which Model Actually Ran

  9. Ashita Orbis 11 min

    Harness Replication: Method, Thirteen Cells and the Artefacts Caught

    The counting proxy, the twelve-task set and its two guards, all thirteen cells, the paired version deltas, the runaway tail, the cross-harness cost table, and the four measurement artefacts that nearly shipped as findings.

    Companion to: What Replicated, and What Did Not

  10. Ashita Orbis 7 min

    DeepSeek V4 Flash against GPT-5.6 Luna: Methodology, Results and Retraction

    Pre-registration, five arms, the full results table, the price sheet, the mislabelled-arm retraction and the build-identity test that established it, plus limitations and spend.

    Companion to: Cheap Tokens, Expensive Answers

  11. Ashita Orbis 8 min

    Cache-Read Coefficient: Methodology and Raw Data

    Full method, instrument notes, phase table, arithmetic, error bars, and limitations for the cache-read coefficient measurement reported in the parent post.

    Companion to: What a Cached Token Actually Costs

  12. Ashita Orbis 8 min

    Context Contamination: How a Coding Harness De-Blinded a Blind Evaluation

    The audit behind the convergent-personality paper's blindness caveat: every evaluative claude -p call in the project ran with the subject's ground-truth trait scores silently in context, because the CLI boots a full agent harness that loads the operator's memory file. Mechanism, empirical canary, blast radius, and the isolation protocol that closed it.

  13. Ashita Orbis 22 min

    Response Profiles Under a Fixed Test: Methods and Full Statistics

    Complete methods, validity tables, response-style diagnostics, and forensic postmortem for the eleven-model personality-battery study: what was measured, what broke in June, and what the remediated pipeline can and cannot support.

    Companion to: How AI Models Describe Themselves Under a Fixed Test

  14. Ashita Orbis 6 min

    The Token Ledger: Generation vs Verification

    The companion post says one day of generation bought four days of verification. Measured in tokens, the imbalance is closer to twenty to one. This is the accounting behind that number: per-message output tokens across a week of building and reviewing the frontier-inference-margins project, with the sensitivity analysis and the caveats a pedantic reader should raise.

    Companion to: Vibe Researching

  15. Ashita Orbis 7 min

    Seven Ghostwriters: The Measurement Data

    The full audit trail for the seven-model style-contract test: access paths, the stylometric measurement table, deterministic checker outcomes, the blind listening protocol, per-letter scores and tells with the unblinded mapping, rank-versus-metrics analysis, known confounds, and the harness difficulty log.

    Companion to: Seven Ghostwriters, One Contract

  16. Ashita Orbis 24 min

    BullshitBench v2: A Methodology Critique

    Statistical methodology critique and extended analysis of BullshitBench v2's rubric design, judge reliability, and scoring philosophy.

    Companion to: Benchmarking "Bullshit Detection"

Same content, three builds · the raw tier is the machine-readable one