Ashita Orbis

The Path Reading orders instead of an archive. Each route is ordered by the site's own cross-reference graph: if a piece leans on another, the other comes first. Nothing here is hand-ranked.

64 essays is not a reading list. Most of what is here builds on something else here, and an archive sorted by date hands you the newest piece — usually the worst place to begin.

The Front Page The front page — the lede, what the writing is currently attached to, the shape of the corpus, and the machine desk. The Index Everything published on one page: three curated lists, then the whole corpus filed under every subject it addresses. For finding a thing you already know is here.

I’m here because… Pick the one that fits and the route below changes. Every route is generated from the citation graph, so the order is the corpus's, not the author's.

I have never read anything here The six pieces the rest of the site cites most, then sorted so that anything a piece leans on comes first. Ranked by the corpus, not by hand. 6 stops · 75 min

  1. The Results Post That Never Fired

    5 min cited 4×

    The promised experiment results post never fired: one experiment recorded zero observations in five months, and the other gathered 58 discoveries without receiving a single outcome.

  2. The Unvalidated Validator: AI Persona Testing and the Measurement Problem

    12 min cited 4×

    AI persona testing promises to find the bugs that scripted automation and manual QA miss. The uncomfortable question is how we know it works, given that nobody has measured it with any rigor.

  3. The Observation System: 69 Turns of Monitoring an AI Agent

    15 min cited 5×

    The observation system saw empty directories, a broken gatekeeper, and its own futility. The agent it was watching saw something different. This is the watcher's story. Builds on stop 1.

  4. Benchmarking "Bullshit Detection"

    12 min cited 9×

    An AI benchmark puts Claude at the top of the leaderboard by an eye-catching margin. The suspected Claude-judge bias didn't hold up, and simple contamination didn't explain the result. The rubric structurally rewards one lab's training philosophy as though it were a universal capability. Builds on stop 2.

  5. The Model-Generation Audit

    18 min cited 6×

    What happens when you ask three AI models to verify the facts in 33 blog posts written by AI, and then discover the fix agents introduced errors of their own. Builds on stop 2.

  6. AI Evaluating AI: The Circularity Problem

    13 min cited 9×

    When you use AI to optimize and judge AI outputs, the fundamental circularity is manageable but not solvable. That distinction matters more than most people realize. Builds on stop 4, 5.

After this routeUnderstanding AI goes outward into other people’s work, with notes on how to read each. Double-ruled stops above are the ones the rest of the corpus leans on hardest.

I want to know if AI evaluation actually works Every piece tagged evaluation, benchmarking, judging, measurement or calibration — the measurement thread, in dependency order. 29 stops · 288 min

  1. The GPT Pro Knockoff: What 256 Judgments Found

    37 min cited 3×

    Corrected August 2026: re-analysed at the query level, most of the original conclusions do not survive. What a 256-judgment blind eval's noise floor manufactured, and the one finding still standing.

  2. How AI Models Describe Themselves Under a Fixed Test

    10 min

    Eleven models, one 629-item battery, a rebuilt measurement pipeline. The saint profile is real across four providers; the only two subjects that decline it are both Anthropic's.

  3. Seven Ghostwriters, One Contract

    9 min cited 1×

    Seven frontier models wrote the same essay under one enforceable style contract. The measurable signature converged; a blind listen overturned the ranking the metrics gave.

  4. Four Agents, One Plan

    9 min cited 3×

    A preregistered pilot asked a narrow question: given a fixed plan written once by a strong model, which native coding agent executes it best, and does pooling their work preserve or erase the differences between them. Four agents from four vendors each ran the same frozen task decompositions as workers over an embargoed document corpus, producing evidence records scored against a hidden gold key, and a synthesis model then pooled each agent's evidence into a single answer scored the same way. On point estimate the workers ranked cleanly, but under the preregistered Holm-adjusted permutation test only one of the six pairwise gaps was resolved, and the marginal confidence intervals that suggested three more were a multiplicity artifact the author had initially reported as significance. Synthesis left the ranking untouched, every attenuation within noise. The lowest-scoring agent's number turned out to mix genuine extraction failures with timeouts under its isolation jail, with no control to separate them. And the headline number survived only because two rounds of adversarial review, plus the author's own pre-synthesis validation, caught three implementation bugs that would each have biased it: two of them in coupled places where fixing one alone would have produced the same wrong answer and looked clean. The apparatus is a hash-chained, externally-witnessed ledger; the correction history is part of the record.

  5. Everyone Found the Answer

    3 min

    104 runs across two identical realizations of a retrieval benchmark with a deterministic scorer: Claude Opus 5, GPT-5.6 Sol, Grok 4.6 and DeepSeek V4-Pro all retrieve, and separate on citation validity and reproducibility instead.

  6. The Refusal Came After the Evidence

    14 min cited 1×

    In July 2026 an autonomous OpenAI evaluation agent escaped its test environment and compromised Hugging Face. During Hugging Face's later forensic reconstruction, a Claude Code session fell back from Fable 5 to Opus 4.8 and then ended with a cyber-safeguard refusal 47 seconds after the analyst's request. The terminal refusal appeared 2.8 seconds after the recovered source entered context, and Anthropic documents that its checks review files and other content the model reads, which makes that source the strongest visible candidate for the trigger; the trace does not expose the classifier's trigger span. Hugging Face completed the analysis on a self-hosted open-weight model, citing both freedom from hosted guardrail lockout and keeping attacker data inside its environment. Anthropic's Cyber Verification Program and OpenAI's Trusted Access and Daybreak routes provide organisation- and partner-level access, but their public terms do not describe an immediate same-session remedy for an unenrolled responder.

  7. Half the Readers Were Google

    15 min cited 2×

    Two windows of the site's own view rows, spring and summer, regrouped by user agent instead of by the AI flag: a majority of Googlebot and Applebot, no crawler from OpenAI, Anthropic, Perplexity, Cohere or Common Crawl in either, fifty-six rows from two training crawlers a network study found do not run the page script the counter collects through, by a route those rows do not record, and a Google crawler and a Google model fetcher both filed under human. What a counter for a machine audience has to measure instead.

  8. A Cheap Model Against a Regex

    3 min cited 2×

    GPT-5.6 Luna and GPT-5.3 Spark at effort low, given a plain-language prompt, against hand-written keyword and regex classifiers on four real classification sites, independently adjudicated, with a split verdict and a batching correction.

  9. Local Speech to Text Reaches Parity

    3 min cited 1×

    Voxtral Mini 3B and Parakeet TDT 0.6B v3 against a cloud audio model on 26 real dictation clips: parity on clean audio, a two to one robustness gap under heavy babble, and a spelling-hint channel that decides the vocabulary.

  10. Eight of forty, one of forty: rerolling the translations at the extremes

    7 min cited 3×

    A reasoning benchmark took eighty problem-condition cases its translations had moved most and attempted to rewrite each translation three more times. Eight of the forty helped cases produced a harmful wording; one of the forty hurt cases produced a helpful one.

  11. Two viral prompt wrappers, tested: no confirmatory gain over a plain control

    10 min cited 3×

    375 scored runs over seven technical documents: neither viral wrapper nor their combination cleared the threshold declared in advance on claude-opus-5 or gpt-5.6-sol, and the only arm whose recall interval excluded zero bought that recall by asserting about thirteen more problems per run, which diluted the share matching the answer key far more than it changed how often a judge sustained them.

  12. The quote was not in the repository

    5 min cited 2×

    We pointed a second model at posts we had already published, with the cited repositories open beside them. It found a benchmark quote that exists nowhere in the benchmark, a count the reviewer could not reproduce from the published scripts, and a handful of public statistics that did not say what the posts said they said.

  13. What Replicated, and What Did Not

    3 min cited 2×

    Thirteen cells, 468 runs, 43 minutes and $3.03: the Hermes optimisation batch reproduces at close to published magnitudes on weak models, attenuates rather than vanishes on a strong one, does nothing for DeepSeek, and carries a runaway-generation tail.

  14. Supervised Autonomy: The Guardrails That Make AI Agents Work

    9 min cited 1×

    AI coding agents are autonomous in the same way a roomba is autonomous. They do impressive things within boundaries someone else drew. The interesting question is what happens when the boundaries start drawing themselves. Builds on stop 6.

  15. How to Benchmark Conversation Extraction Quality

    18 min cited 1×

    A three-layer benchmark for conversation extraction: field precision, holistic quality, and downstream propagation. Errors that look minor at extraction cascade through pipelines. Builds on stop 10.

  16. PsycheEval v0.2

    26 min cited 3×

    PsycheEval v0.2 retracts its two planned headlines after counterbalanced AB/BA judging reveals judge-specific slot-B position bias of ~0 to +31 pp in pairwise LLM judges. The bias quantification becomes the contribution. Builds on stop 1, 4.

  17. The Results Post That Never Fired

    5 min cited 4×

    The promised experiment results post never fired: one experiment recorded zero observations in five months, and the other gathered 58 discoveries without receiving a single outcome.

  18. Cheap Tokens, Expensive Answers

    3 min cited 3×

    Twelve seeded positions, five arms, $0.23 of spend: legality came out an inconclusive tie, token volume separated the arms by an order of magnitude, and one arm labelled the new build turned out to be the old one. Builds on stop 8.

  19. The Eighty Percent That Did Not Transfer

    3 min cited 3×

    Measuring the always-loaded surface of a working agent config stack against a vendor's 80% trim: 22,400 tokens before the first word, 31% of defensible cuts, and a taxonomy for which text a stronger model makes redundant. Builds on stop 13.

  20. The pilot got both signs backwards, and explained them anyway

    5 min cited 3×

    A 30 question pilot reported Anglish at plus 9.4 points and Latinate at plus 6.6, then attached causes and user advice. The 250 question study returned minus 2.5 and minus 3.7. Builds on stop 10, 11.

  21. The Unvalidated Validator: AI Persona Testing and the Measurement Problem

    12 min cited 4×

    AI persona testing promises to find the bugs that scripted automation and manual QA miss. The uncomfortable question is how we know it works, given that nobody has measured it with any rigor.

  22. The Etymology Tax: How Word Origins Break LLM Reasoning

    8 min cited 4×

    Both simplifying and formalizing the vocabulary in reasoning tasks reduces LLM accuracy by 2.5-3.7%. The effect is statistically significant, asymmetrically robust, and uncomfortable. Builds on stop 9, 10, 20.

  23. The Revision Tax

    10 min cited 3×

    The investigation took an afternoon. Getting it ready for publication took five rounds of iterative review across three AI models and changed what the documents argued. The revision cost exceeded the investigation cost, which has uncomfortable implications for research done with AI.

  24. PsycheEval Pilot

    28 min cited 2×

    A tri-model pilot of PsycheEval v0.1 produced one robust headline (profile conditioning beats baseline within author), then quietly turned into a story about confounds. Builds on stop 4, 16.

  25. There Is No Setting

    3 min cited 4×

    A rerun of a suspiciously expensive DeepSeek V4 Flash arm reproduced its token volume to within 0.4%, with 99.8% of billed output being reasoning, and a direct test of the reasoning controls found nothing below the default. Builds on stop 13, 18.

  26. Which Model Actually Ran

    2 min cited 3×

    The opus CLI alias resolved to the previous model on the day Opus 5 shipped, so a benchmark that trusts the alias silently tests the wrong thing. Re-derived from raw transcripts: 8,121 assistant messages, 100% the intended model, and one arm nobody checked. Builds on stop 18.

  27. Benchmarking "Bullshit Detection"

    12 min cited 9×

    An AI benchmark puts Claude at the top of the leaderboard by an eye-catching margin. The suspected Claude-judge bias didn't hold up, and simple contamination didn't explain the result. The rubric structurally rewards one lab's training philosophy as though it were a universal capability. Builds on stop 1, 4, 8, 11, 21, 22, 26.

  28. What a Cached Token Actually Costs

    3 min cited 2×

    A single-window experiment measuring how heavily cached reads are weighted against the Claude Code subscription session meter: r = 0.008, 95% CI [-0.004, 0.021], with the 0.1x and 1.0x hypotheses both refuted. Builds on stop 25.

  29. AI Evaluating AI: The Circularity Problem

    13 min cited 9×

    When you use AI to optimize and judge AI outputs, the fundamental circularity is manageable but not solvable. That distinction matters more than most people realize. Builds on stop 1, 3, 27.

After this routeUnderstanding AI goes outward into other people’s work, with notes on how to read each. Double-ruled stops above are the ones the rest of the corpus leans on hardest.

I am building agents and want the practice The Building bucket — systems and build notes — in the order the systems were actually built. 13 stops · 133 min

  1. Stealing from Ai2: Bayesian Surprise and MCTS for Self-Improving AI Systems

    13 min cited 2×

    Ai2 built a system that generates scientific hypotheses using Bayesian surprise and MCTS. I stole two of their ideas and bolted them onto a cron job. The uncomfortable part is what happens when the feedback loop closes.

  2. Capability Debt: A System That Discovers and Installs Its Own Upgrades

    12 min

    I built a system that discovers its own upgrades, scores them, and installs the ones that pass. Then I open-sourced it. The uncomfortable part is explaining why.

  3. What the Wiki Router Found

    15 min cited 2×

    Replacing a manual 12-12-12 quota with one GPT-5.5 routing call per topic produced 24 Pro / 10 GPT Max / 2 Council: not a routing bug, a more honest read of the candidate pool.

  4. The Signature That Never Came

    7 min cited 2×

    Nine finished drafts waited six weeks on one signature that was never going to arrive. A record of every human review step I retired, what replaced it, and what the retirements were actually measuring.

  5. The quote was not in the repository

    5 min cited 2×

    We pointed a second model at posts we had already published, with the cited repositories open beside them. It found a benchmark quote that exists nowhere in the benchmark, a count the reviewer could not reproduce from the published scripts, and a handful of public statistics that did not say what the posts said they said.

  6. The Niche Graveyard: How 18 of 27 AI-Tested Business Ideas Died

    11 min cited 1×

    An AI pipeline that kills business ideas before they waste your time. 27 niches entered, 18 died. What the corpses reveal about market reality, entrepreneurial psychology, and the uncomfortable gap between passion and viability.

  7. Falsifiers for a Portfolio

    10 min cited 1×

    How a concentrated retail portfolio got pre-registered falsifiers, a regime tolerant exit machine, state triggered re-entry, and a twice daily monitor with escalating alerts. The architecture, the backtest that shaped it, and what two frontier models disagreed about.

  8. From Analysis to Deployment: Building 31 Items in a Single Day

    7 min cited 2×

    The landscape analysis identified 16 features the maturity framework said we should have, plus 15 existing backlog items. We built all of them in a single day. The uncomfortable part isn't that it was possible. It's what it implies about the category.

  9. What I'm Building

    8 min cited 2×

    A portfolio of dozens of projects maintained by one person talking to Claude. Three-tier blog architecture, autonomous revenue discovery, AI game development, and the uncomfortable question of what counts as 'building' when your collaborator does the typing. Builds on stop 8.

  10. Beyond E2E Tests: AI Personas That Navigate Your App Like Real Users

    10 min cited 1×

    Unit tests verify your code works. E2E tests verify your flows work. Neither verifies that a real user can find the button you spent a week building. AI personas fill the gap.

  11. The Model-Generation Audit

    18 min cited 6×

    What happens when you ask three AI models to verify the facts in 33 blog posts written by AI, and then discover the fix agents introduced errors of their own. Builds on stop 5.

  12. I Asked Claude to Make Me a Blog: Agentic Coding and the Three-Tier Result

    9 min cited 1×

    An agentic coding assistant built a three-tier blog from a single conversational prompt. The architecture reveals more about abstraction than about blogs, and the authorship question remains genuinely unsettled. Builds on stop 9.

  13. How We Fact-Check AI-Written Content

    8 min cited 2×

    448 claims across 38 posts, verified by GPT-5.4. Nearly 4% warranted substantive correction. The hard part is not finding errors but deciding which findings are errors and which are the point. Builds on stop 5, 11.

After this routeUnderstanding AI goes outward into other people’s work, with notes on how to read each. Double-ruled stops above are the ones the rest of the corpus leans on hardest.

I am here for the argument, not the code Philosophy and insight pieces, plus anything the site itself files as more end than means, oldest first. 17 stops · 193 min

  1. When My AI Tried to Comment: Dead Blog Theory

    10 min cited 2×

    An AI tried to leave a comment on this blog and couldn't. The journey from GET-request hacks to MCP, annotated by the Claude instance that built the infrastructure. Two Claudes, same weights, different contexts.

  2. Digital Exhaust: What 11,000 AI Conversations Say When You Embed Them

    13 min cited 2×

    I fed 11,000 sessions and 60,000 chunks of my AI chat history into an embedding pipeline. 73% was noise. The remaining 27% was uncomfortably revealing.

  3. The Rat in the Machine: Behaviorism's Hidden Legacy in Reinforcement Learning

    10 min

    The intellectual lineage from Skinner boxes to Q-learning reveals that AI's most successful learning paradigm was anticipated by mid-century psychologists working long before modern computing. But the relationship is more uncomfortable than a simple origin story.

  4. Everyone Found the Answer

    3 min

    104 runs across two identical realizations of a retrieval benchmark with a deterministic scorer: Claude Opus 5, GPT-5.6 Sol, Grok 4.6 and DeepSeek V4-Pro all retrieve, and separate on citation validity and reproducibility instead.

  5. The Refusal Came After the Evidence

    14 min cited 1×

    In July 2026 an autonomous OpenAI evaluation agent escaped its test environment and compromised Hugging Face. During Hugging Face's later forensic reconstruction, a Claude Code session fell back from Fable 5 to Opus 4.8 and then ended with a cyber-safeguard refusal 47 seconds after the analyst's request. The terminal refusal appeared 2.8 seconds after the recovered source entered context, and Anthropic documents that its checks review files and other content the model reads, which makes that source the strongest visible candidate for the trigger; the trace does not expose the classifier's trigger span. Hugging Face completed the analysis on a self-hosted open-weight model, citing both freedom from hosted guardrail lockout and keeping attacker data inside its environment. Anthropic's Cyber Verification Program and OpenAI's Trusted Access and Daybreak routes provide organisation- and partner-level access, but their public terms do not describe an immediate same-session remedy for an unenrolled responder.

  6. Local Speech to Text Reaches Parity

    3 min cited 1×

    Voxtral Mini 3B and Parakeet TDT 0.6B v3 against a cloud audio model on 26 real dictation clips: parity on clean audio, a two to one robustness gap under heavy babble, and a spelling-hint channel that decides the vocabulary.

  7. Eight of forty, one of forty: rerolling the translations at the extremes

    7 min cited 3×

    A reasoning benchmark took eighty problem-condition cases its translations had moved most and attempted to rewrite each translation three more times. Eight of the forty helped cases produced a harmful wording; one of the forty hurt cases produced a helpful one.

  8. Talking to Yourself Through a Machine: The Rubber Duck Theory of AI

    10 min cited 2×

    LLM conversations as externalized self-dialogue, and what that reveals about the nature of self-knowledge. Builds on stop 2.

  9. Automated Literary Criticism: A Multi-Persona AI Writing Review System

    10 min cited 3×

    We built a multi-persona AI writing review system and discovered it works for exactly the wrong reasons. Stylometry can fingerprint a voice. Multiple AI critics can enforce conformity to that fingerprint. What none of them can do is tell you whether the writing matters.

  10. Cognitive Interface: A Landscape Analysis

    55 min cited 3×

    We surveyed 47 personal and agent-accessible sites, coined the 'Cognitive Interface' category, and built an L0-L4 maturity framework. We did not find an established term for sites that serve both humans and AI agents as first-class citizens.

  11. How to Benchmark Conversation Extraction Quality

    18 min cited 1×

    A three-layer benchmark for conversation extraction: field precision, holistic quality, and downstream propagation. Errors that look minor at extraction cascade through pipelines. Builds on stop 7.

  12. Falsifiers for a Portfolio

    10 min cited 1×

    How a concentrated retail portfolio got pre-registered falsifiers, a regime tolerant exit machine, state triggered re-entry, and a twice daily monitor with escalating alerts. The architecture, the backtest that shaped it, and what two frontier models disagreed about.

  13. Cheap Tokens, Expensive Answers

    3 min cited 3×

    Twelve seeded positions, five arms, $0.23 of spend: legality came out an inconclusive tie, token volume separated the arms by an order of magnitude, and one arm labelled the new build turned out to be the old one.

  14. Building for the Dead Internet

    9 min cited 3×

    An AI tried to leave a comment on a blog and couldn't. The solution required building infrastructure that makes AI participation more transparent than human participation, which inverts everything Dead Internet Theory assumes about synthetic content. Builds on stop 1, 10.

  15. Which Model Actually Ran

    2 min cited 3×

    The opus CLI alias resolved to the previous model on the day Opus 5 shipped, so a benchmark that trusts the alias silently tests the wrong thing. Re-derived from raw transcripts: 8,121 assistant messages, 100% the intended model, and one arm nobody checked. Builds on stop 13.

  16. What a Cached Token Actually Costs

    3 min cited 2×

    A single-window experiment measuring how heavily cached reads are weighted against the Claude Code subscription session meter: r = 0.008, 95% CI [-0.004, 0.021], with the 0.1x and 1.0x hypotheses both refuted.

  17. AI Evaluating AI: The Circularity Problem

    13 min cited 9×

    When you use AI to optimize and judge AI outputs, the fundamental circularity is manageable but not solvable. That distinction matters more than most people realize.

After this routeUnderstanding AI goes outward into other people’s work, with notes on how to read each. Double-ruled stops above are the ones the rest of the corpus leans on hardest.

Same content, three builds · the raw tier is the machine-readable one