Ashita Orbis

The Lab Notes

Small experiments on AI models and the tools that run them — findings in the notes, most with a companion investigation holding the method and raw data, published as a starting point for further research.

12 notes, most recent first

  1. 10 min

    Two viral prompt wrappers, tested: no confirmatory gain over a plain control

    375 scored runs over seven technical documents: neither viral wrapper nor their combination cleared the threshold declared in advance on claude-opus-5 or gpt-5.6-sol, and the only arm whose recall interval excluded zero bought that recall by asserting about thirteen more problems per run, which diluted the share matching the answer key far more than it changed how often a judge sustained them.

    Method and raw data: Two Viral Prompt Wrappers: Method, Arms and the Full Grid

  2. 3 min

    Everyone Found the Answer

    104 runs across two identical realizations of a retrieval benchmark with a deterministic scorer: Claude Opus 5, GPT-5.6 Sol, Grok 4.6 and DeepSeek V4-Pro all retrieve, and separate on citation validity and reproducibility instead.

    Method and raw data: Four Models on a Retrieval Contract: Method, Both Realizations, and What the Instrument Cannot Say

  3. 5 min

    The pilot got both signs backwards, and explained them anyway

    A 30 question pilot reported Anglish at plus 9.4 points and Latinate at plus 6.6, then attached causes and user advice. The 250 question study returned minus 2.5 and minus 3.7.

  4. 5 min

    Thirty Fresh Writers and an Empty Log

    An experiment whose entire premise is a controlled information boundary shipped a finished manuscript and left the one instrument that would measure the boundary with no records in it.

  5. 2 min

    Which Model Actually Ran

    The opus CLI alias resolved to the previous model on the day Opus 5 shipped, so a benchmark that trusts the alias silently tests the wrong thing. Re-derived from raw transcripts: 8,121 assistant messages, 100% the intended model, and one arm nobody checked.

    Method and raw data: Model Identity Verification: Methodology and Raw Data

  6. 3 min

    What Replicated, and What Did Not

    Thirteen cells, 468 runs, 43 minutes and $3.03: the Hermes optimisation batch reproduces at close to published magnitudes on weak models, attenuates rather than vanishes on a strong one, does nothing for DeepSeek, and carries a runaway-generation tail.

    Method and raw data: Harness Replication: Method, Thirteen Cells and the Artefacts Caught

  7. 3 min

    There Is No Setting

    A rerun of a suspiciously expensive DeepSeek V4 Flash arm reproduced its token volume to within 0.4%, with 99.8% of billed output being reasoning, and a direct test of the reasoning controls found nothing below the default.

    Method and raw data: The Reasoning Floor: Methodology, Cell Data and Control Probe

  8. 3 min

    The Eighty Percent That Did Not Transfer

    Measuring the always-loaded surface of a working agent config stack against a vendor's 80% trim: 22,400 tokens before the first word, 31% of defensible cuts, and a taxonomy for which text a stronger model makes redundant.

    Method and raw data: Prompt Stack Trim: Method, Taxonomy and the Ranked Proposal

  9. 3 min

    Local Speech to Text Reaches Parity

    Voxtral Mini 3B and Parakeet TDT 0.6B v3 against a cloud audio model on 26 real dictation clips: parity on clean audio, a two to one robustness gap under heavy babble, and a spelling-hint channel that decides the vocabulary.

    Method and raw data: Speech to Text Bake-off: Methodology, Reference Design and Full Tables

  10. 3 min

    Cheap Tokens, Expensive Answers

    Twelve seeded positions, five arms, $0.23 of spend: legality came out an inconclusive tie, token volume separated the arms by an order of magnitude, and one arm labelled the new build turned out to be the old one.

    Method and raw data: DeepSeek V4 Flash against GPT-5.6 Luna: Methodology, Results and Retraction

  11. 3 min

    A Cheap Model Against a Regex

    GPT-5.6 Luna and GPT-5.3 Spark at effort low, given a plain-language prompt, against hand-written keyword and regex classifiers on four real classification sites, independently adjudicated, with a split verdict and a batching correction.

    Method and raw data: Semantic against Keyword: Method, Adjudication and Per-site Verdicts

  12. 3 min

    What a Cached Token Actually Costs

    A single-window experiment measuring how heavily cached reads are weighted against the Claude Code subscription session meter: r = 0.008, 95% CI [-0.004, 0.021], with the 0.1x and 1.0x hypotheses both refuted.

    Method and raw data: Cache-Read Coefficient: Methodology and Raw Data

Same content, three builds · the raw tier is the machine-readable one