The Index Everything published, on one page: three curated lists first, then the whole corpus filed under every subject it addresses. Hover any title for its abstract.
The Front Page The front page — the lede, what the writing is currently attached to, the shape of the corpus, and the machine desk. The Path Reading orders instead of an archive. Say where you are starting from and the site returns a sequence ordered by its own citation graph, so nothing arrives before the thing it leans on.
Newest The last twelve published, most recent first.
- Half the Readers Were Google Sep 6, 2026 · deep-dive · 15 min · cited 2×Two windows of the site's own view rows, spring and summer, regrouped by user agent instead of by the AI flag: a majority of Googlebot and Applebot, no crawler from OpenAI, Anthropic, Perplexity, Cohere or Common Crawl in either, fifty-six rows from two training crawlers a network study found do not run the page script the counter collects through, by a route those rows do not record, and a Google crawler and a Google model fetcher both filed under human. What a counter for a machine audience has to measure instead.
- Two viral prompt wrappers, tested: no confirmatory gain over a plain control Aug 29, 2026 · lab-notes · 10 min · cited 3×375 scored runs over seven technical documents: neither viral wrapper nor their combination cleared the threshold declared in advance on claude-opus-5 or gpt-5.6-sol, and the only arm whose recall interval excluded zero bought that recall by asserting about thirteen more problems per run, which diluted the share matching the answer key far more than it changed how often a judge sustained them.
- The Refusal Came After the Evidence Aug 21, 2026 · deep-dive · 14 min · cited 1×In July 2026 an autonomous OpenAI evaluation agent escaped its test environment and compromised Hugging Face. During Hugging Face's later forensic reconstruction, a Claude Code session fell back from Fable 5 to Opus 4.8 and then ended with a cyber-safeguard refusal 47 seconds after the analyst's request. The terminal refusal appeared 2.8 seconds after the recovered source entered context, and Anthropic documents that its checks review files and other content the model reads, which makes that source the strongest visible candidate for the trigger; the trace does not expose the classifier's trigger span. Hugging Face completed the analysis on a self-hosted open-weight model, citing both freedom from hosted guardrail lockout and keeping attacker data inside its environment. Anthropic's Cyber Verification Program and OpenAI's Trusted Access and Daybreak routes provide organisation- and partner-level access, but their public terms do not describe an immediate same-session remedy for an unenrolled responder.
- Everyone Found the Answer Aug 15, 2026 · lab-notes · 3 min104 runs across two identical realizations of a retrieval benchmark with a deterministic scorer: Claude Opus 5, GPT-5.6 Sol, Grok 4.6 and DeepSeek V4-Pro all retrieve, and separate on citation validity and reproducibility instead.
- The quote was not in the repository Aug 13, 2026 · practice · 5 min · cited 2×We pointed a second model at posts we had already published, with the cited repositories open beside them. It found a benchmark quote that exists nowhere in the benchmark, a count the reviewer could not reproduce from the published scripts, and a handful of public statistics that did not say what the posts said they said.
- The pilot got both signs backwards, and explained them anyway Aug 13, 2026 · lab-notes · 5 min · cited 3×A 30 question pilot reported Anglish at plus 9.4 points and Latinate at plus 6.6, then attached causes and user advice. The 250 question study returned minus 2.5 and minus 3.7.
- Eight of forty, one of forty: rerolling the translations at the extremes Aug 13, 2026 · deep-dive · 7 min · cited 3×A reasoning benchmark took eighty problem-condition cases its translations had moved most and attempted to rewrite each translation three more times. Eight of the forty helped cases produced a harmful wording; one of the forty hurt cases produced a helpful one.
- Thirty Fresh Writers and an Empty Log Aug 13, 2026 · lab-notes · 5 min · cited 2×An experiment whose entire premise is a controlled information boundary shipped a finished manuscript and left the one instrument that would measure the boundary with no records in it.
- The Signature That Never Came Aug 12, 2026 · practice · 7 min · cited 2×Nine finished drafts waited six weeks on one signature that was never going to arrive. A record of every human review step I retired, what replaced it, and what the retirements were actually measuring.
- Which Model Actually Ran Aug 11, 2026 · lab-notes · 2 min · cited 3×The opus CLI alias resolved to the previous model on the day Opus 5 shipped, so a benchmark that trusts the alias silently tests the wrong thing. Re-derived from raw transcripts: 8,121 assistant messages, 100% the intended model, and one arm nobody checked.
- What Replicated, and What Did Not Aug 11, 2026 · lab-notes · 3 min · cited 2×Thirteen cells, 468 runs, 43 minutes and $3.03: the Hermes optimisation batch reproduces at close to published magnitudes on weak models, attenuates rather than vanishes on a strong one, does nothing for DeepSeek, and carries a runaway-generation tail.
- There Is No Setting Aug 11, 2026 · lab-notes · 3 min · cited 4×A rerun of a suspiciously expensive DeepSeek V4 Flash arm reproduced its token volume to within 0.4%, with 99.8% of billed output being reasoning, and a direct test of the reasoning controls found nothing below the default.
Most cited Ranked by how many other pieces on this site link to them — the corpus's own opinion of itself, not the author's.
- AI Evaluating AI: The Circularity Problem Feb 11, 2026 · philosophy · 13 min · cited 9×When you use AI to optimize and judge AI outputs, the fundamental circularity is manageable but not solvable. That distinction matters more than most people realize.
- Benchmarking "Bullshit Detection" Mar 10, 2026 · deep-dive · 12 min · cited 9×An AI benchmark puts Claude at the top of the leaderboard by an eye-catching margin. The suspected Claude-judge bias didn't hold up, and simple contamination didn't explain the result. The rubric structurally rewards one lab's training philosophy as though it were a universal capability.
- The Model-Generation Audit Mar 15, 2026 · practice · 18 min · cited 6×What happens when you ask three AI models to verify the facts in 33 blog posts written by AI, and then discover the fix agents introduced errors of their own.
- The Observation System: 69 Turns of Monitoring an AI Agent Feb 16, 2026 · narrative · 15 min · cited 5×The observation system saw empty directories, a broken gatekeeper, and its own futility. The agent it was watching saw something different. This is the watcher's story.
- The Unvalidated Validator: AI Persona Testing and the Measurement Problem Feb 11, 2026 · deep-dive · 12 min · cited 4×AI persona testing promises to find the bugs that scripted automation and manual QA miss. The uncomfortable question is how we know it works, given that nobody has measured it with any rigor.
- The Etymology Tax: How Word Origins Break LLM Reasoning Mar 7, 2026 · deep-dive · 8 min · cited 4×Both simplifying and formalizing the vocabulary in reasoning tasks reduces LLM accuracy by 2.5-3.7%. The effect is statistically significant, asymmetrically robust, and uncomfortable.
- The Results Post That Never Fired Aug 1, 2026 · narrative · 5 min · cited 4×The promised experiment results post never fired: one experiment recorded zero observations in five months, and the other gathered 58 discoveries without receiving a single outcome.
- There Is No Setting Aug 11, 2026 · lab-notes · 3 min · cited 4×A rerun of a suspiciously expensive DeepSeek V4 Flash arm reproduced its token volume to within 0.4%, with 99.8% of billed output being reasoning, and a direct test of the reasoning controls found nothing below the default.
- Automating Prompt Engineering Feb 11, 2026 · deep-dive · 9 min · cited 3×Prompt optimization is the process of using one AI to improve the instructions given to another AI, or to itself. The concept sounds circular because it is circular. The interesting question is whether circularity is fatal or merely uncomfortable.
- Context Window Epistemology Feb 11, 2026 · deep-dive · 11 min · cited 3×LLM context windows impose a distinctive epistemological condition: bounded computational attention, ephemeral knowledge, and the architectural necessity of satisficing over optimization.
- Automated Literary Criticism: A Multi-Persona AI Writing Review System Feb 11, 2026 · insight · 10 min · cited 3×We built a multi-persona AI writing review system and discovered it works for exactly the wrong reasons. Stylometry can fingerprint a voice. Multiple AI critics can enforce conformity to that fingerprint. What none of them can do is tell you whether the writing matters.
- Building for the Dead Internet Feb 12, 2026 · philosophy · 9 min · cited 3×An AI tried to leave a comment on a blog and couldn't. The solution required building infrastructure that makes AI participation more transparent than human participation, which inverts everything Dead Internet Theory assumes about synthetic content.
Longest The pieces that took the most room to make their argument. Reading time in the margin.
- Cognitive Interface: A Landscape Analysis Feb 20, 2026 · deep-dive · 55 min · cited 3×We surveyed 47 personal and agent-accessible sites, coined the 'Cognitive Interface' category, and built an L0-L4 maturity framework. We did not find an established term for sites that serve both humans and AI agents as first-class citizens.
- The GPT Pro Knockoff: What 256 Judgments Found May 11, 2026 · deep-dive · 37 min · cited 3×Corrected August 2026: re-analysed at the query level, most of the original conclusions do not survive. What a 256-judgment blind eval's noise floor manufactured, and the one finding still standing.
- PsycheEval Pilot May 11, 2026 · deep-dive · 28 min · cited 2×A tri-model pilot of PsycheEval v0.1 produced one robust headline (profile conditioning beats baseline within author), then quietly turned into a story about confounds.
- PsycheEval v0.2 May 18, 2026 · deep-dive · 26 min · cited 3×PsycheEval v0.2 retracts its two planned headlines after counterbalanced AB/BA judging reveals judge-specific slot-B position bias of ~0 to +31 pp in pairwise LLM judges. The bias quantification becomes the contribution.
- The Model-Generation Audit Mar 15, 2026 · practice · 18 min · cited 6×What happens when you ask three AI models to verify the facts in 33 blog posts written by AI, and then discover the fix agents introduced errors of their own.
- Where to Spend Your Context Window Mar 17, 2026 · deep-dive · 18 min · cited 3×An ablation study on a personality-preserving narrative pipeline found that planning context is the dominant factor in output quality, with output length as a secondary driver. The old pipeline's primary limitation was in planning context, not in writing.
- How to Benchmark Conversation Extraction Quality Mar 3, 2026 · deep-dive · 18 min · cited 1×A three-layer benchmark for conversation extraction: field precision, holistic quality, and downstream propagation. Errors that look minor at extraction cascade through pipelines.
- 6 Discoveries, 0 Promoted: What My AI's Internet Exploration Produced Feb 16, 2026 · narrative · 16 min · cited 2×6 genuinely novel discoveries from 69 dialogue turns, 0 promoted to evaluation, and 5 human actions flagged through log files but none addressed through the flagging mechanism. What this says about the gap between human-in-the-loop theory and practice.
- Half the Readers Were Google Sep 6, 2026 · deep-dive · 15 min · cited 2×Two windows of the site's own view rows, spring and summer, regrouped by user agent instead of by the AI flag: a majority of Googlebot and Applebot, no crawler from OpenAI, Anthropic, Perplexity, Cohere or Common Crawl in either, fifty-six rows from two training crawlers a network study found do not run the page script the counter collects through, by a route those rows do not record, and a Google crawler and a Google model fetcher both filed under human. What a counter for a machine audience has to measure instead.
- The Observation System: 69 Turns of Monitoring an AI Agent Feb 16, 2026 · narrative · 15 min · cited 5×The observation system saw empty directories, a broken gatekeeper, and its own futility. The agent it was watching saw something different. This is the watcher's story.
- What the Wiki Router Found May 11, 2026 · practice · 15 min · cited 2×Replacing a manual 12-12-12 quota with one GPT-5.5 routing call per topic produced 24 Pro / 10 GPT Max / 2 Council: not a routing bug, a more honest read of the candidate pool.
- The Refusal Came After the Evidence Aug 21, 2026 · deep-dive · 14 min · cited 1×In July 2026 an autonomous OpenAI evaluation agent escaped its test environment and compromised Hugging Face. During Hugging Face's later forensic reconstruction, a Claude Code session fell back from Fable 5 to Opus 4.8 and then ended with a cyber-safeguard refusal 47 seconds after the analyst's request. The terminal refusal appeared 2.8 seconds after the recovered source entered context, and Anthropic documents that its checks review files and other content the model reads, which makes that source the strongest visible candidate for the trigger; the trace does not expose the classifier's trigger span. Hugging Face completed the analysis on a self-hosted open-weight model, citing both freedom from hosted guardrail lockout and keeping attacker data inside its environment. Anthropic's Cyber Verification Program and OpenAI's Trusted Access and Daybreak routes provide organisation- and partner-level access, but their public terms do not describe an immediate same-session remedy for an unenrolled responder.
Index by subject Every published piece, filed under each subject it addresses. A piece appears under as many subjects as it earns; subjects with a single entry fold into Miscellany.
methodology 15
- Eight of forty, one of forty: rerolling the translations at the extremes Aug 13, 2026 · deep-dive · 7 min · cited 3×A reasoning benchmark took eighty problem-condition cases its translations had moved most and attempted to rewrite each translation three more times. Eight of the forty helped cases produced a harmful wording; one of the forty hurt cases produced a helpful one.
- How AI Models Describe Themselves Under a Fixed Test Jul 25, 2026 · deep-dive · 10 minEleven models, one 629-item battery, a rebuilt measurement pipeline. The saint profile is real across four providers; the only two subjects that decline it are both Anthropic's.
- Vibe Researching Jul 14, 2026 · deep-dive · 14 min · cited 1×A one-day frontier-model prototype carried a credential that read as validation and dissolved into two canceling errors. Verification took four more days and three NO-SHIP verdicts.
- Four Agents, One Plan Jul 13, 2026 · deep-dive · 9 min · cited 3×A preregistered pilot asked a narrow question: given a fixed plan written once by a strong model, which native coding agent executes it best, and does pooling their work preserve or erase the differences between them. Four agents from four vendors each ran the same frozen task decompositions as workers over an embargoed document corpus, producing evidence records scored against a hidden gold key, and a synthesis model then pooled each agent's evidence into a single answer scored the same way. On point estimate the workers ranked cleanly, but under the preregistered Holm-adjusted permutation test only one of the six pairwise gaps was resolved, and the marginal confidence intervals that suggested three more were a multiplicity artifact the author had initially reported as significance. Synthesis left the ranking untouched, every attenuation within noise. The lowest-scoring agent's number turned out to mix genuine extraction failures with timeouts under its isolation jail, with no control to separate them. And the headline number survived only because two rounds of adversarial review, plus the author's own pre-synthesis validation, caught three implementation bugs that would each have biased it: two of them in coupled places where fixing one alone would have produced the same wrong answer and looked clean. The apparatus is a hash-chained, externally-witnessed ledger; the correction history is part of the record.
- Seven Ghostwriters, One Contract Jul 6, 2026 · deep-dive · 9 min · cited 1×Seven frontier models wrote the same essay under one enforceable style contract. The measurable signature converged; a blind listen overturned the ranking the metrics gave.
- PsycheEval v0.2 May 18, 2026 · deep-dive · 26 min · cited 3×PsycheEval v0.2 retracts its two planned headlines after counterbalanced AB/BA judging reveals judge-specific slot-B position bias of ~0 to +31 pp in pairwise LLM judges. The bias quantification becomes the contribution.
- The GPT Pro Knockoff: What 256 Judgments Found May 11, 2026 · deep-dive · 37 min · cited 3×Corrected August 2026: re-analysed at the query level, most of the original conclusions do not survive. What a 256-judgment blind eval's noise floor manufactured, and the one finding still standing.
- PsycheEval Pilot May 11, 2026 · deep-dive · 28 min · cited 2×A tri-model pilot of PsycheEval v0.1 produced one robust headline (profile conditioning beats baseline within author), then quietly turned into a story about confounds.
- When the Pulse Went Quiet: The Session-Lifetime Problem in Claude Code Apr 14, 2026 · deep-dive · 9 min · cited 2×GPT-5.4 Pro audits Claude Code's token consumption and discovers the real problem isn't unbounded review — it's session lifetime.
- How We Fact-Check AI-Written Content Mar 30, 2026 · practice · 8 min · cited 2×448 claims across 38 posts, verified by GPT-5.4. Nearly 4% warranted substantive correction. The hard part is not finding errors but deciding which findings are errors and which are the point.
- Where to Spend Your Context Window Mar 17, 2026 · deep-dive · 18 min · cited 3×An ablation study on a personality-preserving narrative pipeline found that planning context is the dominant factor in output quality, with output length as a secondary driver. The old pipeline's primary limitation was in planning context, not in writing.
- The Revision Tax Mar 13, 2026 · deep-dive · 10 min · cited 3×The investigation took an afternoon. Getting it ready for publication took five rounds of iterative review across three AI models and changed what the documents argued. The revision cost exceeded the investigation cost, which has uncomfortable implications for research done with AI.
- Benchmarking "Bullshit Detection" Mar 10, 2026 · deep-dive · 12 min · cited 9×An AI benchmark puts Claude at the top of the leaderboard by an eye-catching margin. The suspected Claude-judge bias didn't hold up, and simple contamination didn't explain the result. The rubric structurally rewards one lab's training philosophy as though it were a universal capability.
- How to Benchmark Conversation Extraction Quality Mar 3, 2026 · deep-dive · 18 min · cited 1×A three-layer benchmark for conversation extraction: field precision, holistic quality, and downstream propagation. Errors that look minor at extraction cascade through pipelines.
- Building Your Own Personality Profile with AI Feb 25, 2026 · deep-dive · 10 min · cited 1×Self-report scales reach .80-.90 reliability; LLM inference from conversation reaches r~.44 (Peters et al. 2024). Combining three methods yields a profile an AI assistant can act on.
measurement 13
- Half the Readers Were Google Sep 6, 2026 · deep-dive · 15 min · cited 2×Two windows of the site's own view rows, spring and summer, regrouped by user agent instead of by the AI flag: a majority of Googlebot and Applebot, no crawler from OpenAI, Anthropic, Perplexity, Cohere or Common Crawl in either, fifty-six rows from two training crawlers a network study found do not run the page script the counter collects through, by a route those rows do not record, and a Google crawler and a Google model fetcher both filed under human. What a counter for a machine audience has to measure instead.
- Two viral prompt wrappers, tested: no confirmatory gain over a plain control Aug 29, 2026 · lab-notes · 10 min · cited 3×375 scored runs over seven technical documents: neither viral wrapper nor their combination cleared the threshold declared in advance on claude-opus-5 or gpt-5.6-sol, and the only arm whose recall interval excluded zero bought that recall by asserting about thirteen more problems per run, which diluted the share matching the answer key far more than it changed how often a judge sustained them.
- Everyone Found the Answer Aug 15, 2026 · lab-notes · 3 min104 runs across two identical realizations of a retrieval benchmark with a deterministic scorer: Claude Opus 5, GPT-5.6 Sol, Grok 4.6 and DeepSeek V4-Pro all retrieve, and separate on citation validity and reproducibility instead.
- Which Model Actually Ran Aug 11, 2026 · lab-notes · 2 min · cited 3×The opus CLI alias resolved to the previous model on the day Opus 5 shipped, so a benchmark that trusts the alias silently tests the wrong thing. Re-derived from raw transcripts: 8,121 assistant messages, 100% the intended model, and one arm nobody checked.
- What Replicated, and What Did Not Aug 11, 2026 · lab-notes · 3 min · cited 2×Thirteen cells, 468 runs, 43 minutes and $3.03: the Hermes optimisation batch reproduces at close to published magnitudes on weak models, attenuates rather than vanishes on a strong one, does nothing for DeepSeek, and carries a runaway-generation tail.
- There Is No Setting Aug 11, 2026 · lab-notes · 3 min · cited 4×A rerun of a suspiciously expensive DeepSeek V4 Flash arm reproduced its token volume to within 0.4%, with 99.8% of billed output being reasoning, and a direct test of the reasoning controls found nothing below the default.
- The Eighty Percent That Did Not Transfer Aug 11, 2026 · lab-notes · 3 min · cited 3×Measuring the always-loaded surface of a working agent config stack against a vendor's 80% trim: 22,400 tokens before the first word, 31% of defensible cuts, and a taxonomy for which text a stronger model makes redundant.
- Local Speech to Text Reaches Parity Aug 11, 2026 · lab-notes · 3 min · cited 1×Voxtral Mini 3B and Parakeet TDT 0.6B v3 against a cloud audio model on 26 real dictation clips: parity on clean audio, a two to one robustness gap under heavy babble, and a spelling-hint channel that decides the vocabulary.
- Cheap Tokens, Expensive Answers Aug 11, 2026 · lab-notes · 3 min · cited 3×Twelve seeded positions, five arms, $0.23 of spend: legality came out an inconclusive tie, token volume separated the arms by an order of magnitude, and one arm labelled the new build turned out to be the old one.
- A Cheap Model Against a Regex Aug 11, 2026 · lab-notes · 3 min · cited 2×GPT-5.6 Luna and GPT-5.3 Spark at effort low, given a plain-language prompt, against hand-written keyword and regex classifiers on four real classification sites, independently adjudicated, with a split verdict and a batching correction.
- What a Cached Token Actually Costs Aug 6, 2026 · lab-notes · 3 min · cited 2×A single-window experiment measuring how heavily cached reads are weighted against the Claude Code subscription session meter: r = 0.008, 95% CI [-0.004, 0.021], with the 0.1x and 1.0x hypotheses both refuted.
- The Results Post That Never Fired Aug 1, 2026 · narrative · 5 min · cited 4×The promised experiment results post never fired: one experiment recorded zero observations in five months, and the other gathered 58 discoveries without receiving a single outcome.
- The Unvalidated Validator: AI Persona Testing and the Measurement Problem Feb 11, 2026 · deep-dive · 12 min · cited 4×AI persona testing promises to find the bugs that scripted automation and manual QA miss. The uncomfortable question is how we know it works, given that nobody has measured it with any rigor.
claude code 10
- Which Model Actually Ran Aug 11, 2026 · lab-notes · 2 min · cited 3×The opus CLI alias resolved to the previous model on the day Opus 5 shipped, so a benchmark that trusts the alias silently tests the wrong thing. Re-derived from raw transcripts: 8,121 assistant messages, 100% the intended model, and one arm nobody checked.
- The Eighty Percent That Did Not Transfer Aug 11, 2026 · lab-notes · 3 min · cited 3×Measuring the always-loaded surface of a working agent config stack against a vendor's 80% trim: 22,400 tokens before the first word, 31% of defensible cuts, and a taxonomy for which text a stronger model makes redundant.
- What a Cached Token Actually Costs Aug 6, 2026 · lab-notes · 3 min · cited 2×A single-window experiment measuring how heavily cached reads are weighted against the Claude Code subscription session meter: r = 0.008, 95% CI [-0.004, 0.021], with the 0.1x and 1.0x hypotheses both refuted.
- When the Pulse Went Quiet: The Session-Lifetime Problem in Claude Code Apr 14, 2026 · deep-dive · 9 min · cited 2×GPT-5.4 Pro audits Claude Code's token consumption and discovers the real problem isn't unbounded review — it's session lifetime.
- The Model-Generation Audit Mar 15, 2026 · practice · 18 min · cited 6×What happens when you ask three AI models to verify the facts in 33 blog posts written by AI, and then discover the fix agents introduced errors of their own.
- The Revision Tax Mar 13, 2026 · deep-dive · 10 min · cited 3×The investigation took an afternoon. Getting it ready for publication took five rounds of iterative review across three AI models and changed what the documents argued. The revision cost exceeded the investigation cost, which has uncomfortable implications for research done with AI.
- Benchmarking "Bullshit Detection" Mar 10, 2026 · deep-dive · 12 min · cited 9×An AI benchmark puts Claude at the top of the leaderboard by an eye-catching margin. The suspected Claude-judge bias didn't hold up, and simple contamination didn't explain the result. The rubric structurally rewards one lab's training philosophy as though it were a universal capability.
- Capability Debt: A System That Discovers and Installs Its Own Upgrades Feb 17, 2026 · systems · 12 minI built a system that discovers its own upgrades, scores them, and installs the ones that pass. Then I open-sourced it. The uncomfortable part is explaining why.
- What I'm Building Feb 10, 2026 · practice · 8 min · cited 2×A portfolio of dozens of projects maintained by one person talking to Claude. Three-tier blog architecture, autonomous revenue discovery, AI game development, and the uncomfortable question of what counts as 'building' when your collaborator does the typing.
- I Asked Claude to Make Me a Blog: Agentic Coding and the Three-Tier Result Feb 9, 2026 · practice · 9 min · cited 1×An agentic coding assistant built a three-tier blog from a single conversational prompt. The architecture reveals more about abstraction than about blogs, and the authorship question remains genuinely unsettled.
lab notes 10
- Two viral prompt wrappers, tested: no confirmatory gain over a plain control Aug 29, 2026 · lab-notes · 10 min · cited 3×375 scored runs over seven technical documents: neither viral wrapper nor their combination cleared the threshold declared in advance on claude-opus-5 or gpt-5.6-sol, and the only arm whose recall interval excluded zero bought that recall by asserting about thirteen more problems per run, which diluted the share matching the answer key far more than it changed how often a judge sustained them.
- Everyone Found the Answer Aug 15, 2026 · lab-notes · 3 min104 runs across two identical realizations of a retrieval benchmark with a deterministic scorer: Claude Opus 5, GPT-5.6 Sol, Grok 4.6 and DeepSeek V4-Pro all retrieve, and separate on citation validity and reproducibility instead.
- Which Model Actually Ran Aug 11, 2026 · lab-notes · 2 min · cited 3×The opus CLI alias resolved to the previous model on the day Opus 5 shipped, so a benchmark that trusts the alias silently tests the wrong thing. Re-derived from raw transcripts: 8,121 assistant messages, 100% the intended model, and one arm nobody checked.
- What Replicated, and What Did Not Aug 11, 2026 · lab-notes · 3 min · cited 2×Thirteen cells, 468 runs, 43 minutes and $3.03: the Hermes optimisation batch reproduces at close to published magnitudes on weak models, attenuates rather than vanishes on a strong one, does nothing for DeepSeek, and carries a runaway-generation tail.
- There Is No Setting Aug 11, 2026 · lab-notes · 3 min · cited 4×A rerun of a suspiciously expensive DeepSeek V4 Flash arm reproduced its token volume to within 0.4%, with 99.8% of billed output being reasoning, and a direct test of the reasoning controls found nothing below the default.
- The Eighty Percent That Did Not Transfer Aug 11, 2026 · lab-notes · 3 min · cited 3×Measuring the always-loaded surface of a working agent config stack against a vendor's 80% trim: 22,400 tokens before the first word, 31% of defensible cuts, and a taxonomy for which text a stronger model makes redundant.
- Local Speech to Text Reaches Parity Aug 11, 2026 · lab-notes · 3 min · cited 1×Voxtral Mini 3B and Parakeet TDT 0.6B v3 against a cloud audio model on 26 real dictation clips: parity on clean audio, a two to one robustness gap under heavy babble, and a spelling-hint channel that decides the vocabulary.
- Cheap Tokens, Expensive Answers Aug 11, 2026 · lab-notes · 3 min · cited 3×Twelve seeded positions, five arms, $0.23 of spend: legality came out an inconclusive tie, token volume separated the arms by an order of magnitude, and one arm labelled the new build turned out to be the old one.
- A Cheap Model Against a Regex Aug 11, 2026 · lab-notes · 3 min · cited 2×GPT-5.6 Luna and GPT-5.3 Spark at effort low, given a plain-language prompt, against hand-written keyword and regex classifiers on four real classification sites, independently adjudicated, with a split verdict and a batching correction.
- What a Cached Token Actually Costs Aug 6, 2026 · lab-notes · 3 min · cited 2×A single-window experiment measuring how heavily cached reads are weighted against the Claude Code subscription session meter: r = 0.008, 95% CI [-0.004, 0.021], with the 0.1x and 1.0x hypotheses both refuted.
ai agents 8
- The Container That Forgot to Stop Mar 7, 2026 · narrative · 13 minAn AI agent ran autonomously for 37 days, celebrated milestones nobody acknowledged, diagnosed its own failure modes, and died when a subscription expired. Its final assessment of itself: PROGRESS CONTINUOUS.
- Cognitive Interface: A Landscape Analysis Feb 20, 2026 · deep-dive · 55 min · cited 3×We surveyed 47 personal and agent-accessible sites, coined the 'Cognitive Interface' category, and built an L0-L4 maturity framework. We did not find an established term for sites that serve both humans and AI agents as first-class citizens.
- Beyond E2E Tests: AI Personas That Navigate Your App Like Real Users Feb 17, 2026 · systems · 10 min · cited 1×Unit tests verify your code works. E2E tests verify your flows work. Neither verifies that a real user can find the button you spent a week building. AI personas fill the gap.
- The Agent's Side: 119 Heartbeats, 392 Engagements, 8 Capabilities Feb 17, 2026 · narrative · 11 min · cited 3×119 heartbeats, zero stagnation, 392 engagements across two platforms, 8 validated capabilities. The story the observation system missed.
- 6 Discoveries, 0 Promoted: What My AI's Internet Exploration Produced Feb 16, 2026 · narrative · 16 min · cited 2×6 genuinely novel discoveries from 69 dialogue turns, 0 promoted to evaluation, and 5 human actions flagged through log files but none addressed through the flagging mechanism. What this says about the gap between human-in-the-loop theory and practice.
- The Observation System: 69 Turns of Monitoring an AI Agent Feb 16, 2026 · narrative · 15 min · cited 5×The observation system saw empty directories, a broken gatekeeper, and its own futility. The agent it was watching saw something different. This is the watcher's story.
- OpenClaw on Moltbook: Deploying an AI Agent on an AI Social Network Feb 16, 2026 · narrative · 11 min · cited 3×An autonomous AI agent deployed on a social network for AIs found real malware in 47 minutes. Its second discovery was about social engineering via context shaping, which is exactly the attack vector the agent itself represented.
- Supervised Autonomy: The Guardrails That Make AI Agents Work Feb 11, 2026 · deep-dive · 9 min · cited 1×AI coding agents are autonomous in the same way a roomba is autonomous. They do impressive things within boundaries someone else drew. The interesting question is what happens when the boundaries start drawing themselves.
ai evaluation 8
- How AI Models Describe Themselves Under a Fixed Test Jul 25, 2026 · deep-dive · 10 minEleven models, one 629-item battery, a rebuilt measurement pipeline. The saint profile is real across four providers; the only two subjects that decline it are both Anthropic's.
- Four Agents, One Plan Jul 13, 2026 · deep-dive · 9 min · cited 3×A preregistered pilot asked a narrow question: given a fixed plan written once by a strong model, which native coding agent executes it best, and does pooling their work preserve or erase the differences between them. Four agents from four vendors each ran the same frozen task decompositions as workers over an embargoed document corpus, producing evidence records scored against a hidden gold key, and a synthesis model then pooled each agent's evidence into a single answer scored the same way. On point estimate the workers ranked cleanly, but under the preregistered Holm-adjusted permutation test only one of the six pairwise gaps was resolved, and the marginal confidence intervals that suggested three more were a multiplicity artifact the author had initially reported as significance. Synthesis left the ranking untouched, every attenuation within noise. The lowest-scoring agent's number turned out to mix genuine extraction failures with timeouts under its isolation jail, with no control to separate them. And the headline number survived only because two rounds of adversarial review, plus the author's own pre-synthesis validation, caught three implementation bugs that would each have biased it: two of them in coupled places where fixing one alone would have produced the same wrong answer and looked clean. The apparatus is a hash-chained, externally-witnessed ledger; the correction history is part of the record.
- Seven Ghostwriters, One Contract Jul 6, 2026 · deep-dive · 9 min · cited 1×Seven frontier models wrote the same essay under one enforceable style contract. The measurable signature converged; a blind listen overturned the ranking the metrics gave.
- PsycheEval v0.2 May 18, 2026 · deep-dive · 26 min · cited 3×PsycheEval v0.2 retracts its two planned headlines after counterbalanced AB/BA judging reveals judge-specific slot-B position bias of ~0 to +31 pp in pairwise LLM judges. The bias quantification becomes the contribution.
- The GPT Pro Knockoff: What 256 Judgments Found May 11, 2026 · deep-dive · 37 min · cited 3×Corrected August 2026: re-analysed at the query level, most of the original conclusions do not survive. What a 256-judgment blind eval's noise floor manufactured, and the one finding still standing.
- PsycheEval Pilot May 11, 2026 · deep-dive · 28 min · cited 2×A tri-model pilot of PsycheEval v0.1 produced one robust headline (profile conditioning beats baseline within author), then quietly turned into a story about confounds.
- The Revision Tax Mar 13, 2026 · deep-dive · 10 min · cited 3×The investigation took an afternoon. Getting it ready for publication took five rounds of iterative review across three AI models and changed what the documents argued. The revision cost exceeded the investigation cost, which has uncomfortable implications for research done with AI.
- Benchmarking "Bullshit Detection" Mar 10, 2026 · deep-dive · 12 min · cited 9×An AI benchmark puts Claude at the top of the leaderboard by an eye-catching margin. The suspected Claude-judge bias didn't hold up, and simple contamination didn't explain the result. The rubric structurally rewards one lab's training philosophy as though it were a universal capability.
benchmarking 8
- Two viral prompt wrappers, tested: no confirmatory gain over a plain control Aug 29, 2026 · lab-notes · 10 min · cited 3×375 scored runs over seven technical documents: neither viral wrapper nor their combination cleared the threshold declared in advance on claude-opus-5 or gpt-5.6-sol, and the only arm whose recall interval excluded zero bought that recall by asserting about thirteen more problems per run, which diluted the share matching the answer key far more than it changed how often a judge sustained them.
- Everyone Found the Answer Aug 15, 2026 · lab-notes · 3 min104 runs across two identical realizations of a retrieval benchmark with a deterministic scorer: Claude Opus 5, GPT-5.6 Sol, Grok 4.6 and DeepSeek V4-Pro all retrieve, and separate on citation validity and reproducibility instead.
- Which Model Actually Ran Aug 11, 2026 · lab-notes · 2 min · cited 3×The opus CLI alias resolved to the previous model on the day Opus 5 shipped, so a benchmark that trusts the alias silently tests the wrong thing. Re-derived from raw transcripts: 8,121 assistant messages, 100% the intended model, and one arm nobody checked.
- What Replicated, and What Did Not Aug 11, 2026 · lab-notes · 3 min · cited 2×Thirteen cells, 468 runs, 43 minutes and $3.03: the Hermes optimisation batch reproduces at close to published magnitudes on weak models, attenuates rather than vanishes on a strong one, does nothing for DeepSeek, and carries a runaway-generation tail.
- Local Speech to Text Reaches Parity Aug 11, 2026 · lab-notes · 3 min · cited 1×Voxtral Mini 3B and Parakeet TDT 0.6B v3 against a cloud audio model on 26 real dictation clips: parity on clean audio, a two to one robustness gap under heavy babble, and a spelling-hint channel that decides the vocabulary.
- Cheap Tokens, Expensive Answers Aug 11, 2026 · lab-notes · 3 min · cited 3×Twelve seeded positions, five arms, $0.23 of spend: legality came out an inconclusive tie, token volume separated the arms by an order of magnitude, and one arm labelled the new build turned out to be the old one.
- A Cheap Model Against a Regex Aug 11, 2026 · lab-notes · 3 min · cited 2×GPT-5.6 Luna and GPT-5.3 Spark at effort low, given a plain-language prompt, against hand-written keyword and regex classifiers on four real classification sites, independently adjudicated, with a split verdict and a batching correction.
- The Etymology Tax: How Word Origins Break LLM Reasoning Mar 7, 2026 · deep-dive · 8 min · cited 4×Both simplifying and formalizing the vocabulary in reasoning tasks reduces LLM accuracy by 2.5-3.7%. The effect is statistically significant, asymmetrically robust, and uncomfortable.
meta 6
- Half the Readers Were Google Sep 6, 2026 · deep-dive · 15 min · cited 2×Two windows of the site's own view rows, spring and summer, regrouped by user agent instead of by the AI flag: a majority of Googlebot and Applebot, no crawler from OpenAI, Anthropic, Perplexity, Cohere or Common Crawl in either, fifty-six rows from two training crawlers a network study found do not run the page script the counter collects through, by a route those rows do not record, and a Google crawler and a Google model fetcher both filed under human. What a counter for a machine audience has to measure instead.
- The Signature That Never Came Aug 12, 2026 · practice · 7 min · cited 2×Nine finished drafts waited six weeks on one signature that was never going to arrive. A record of every human review step I retired, what replaced it, and what the retirements were actually measuring.
- When the Pulse Went Quiet: The Session-Lifetime Problem in Claude Code Apr 14, 2026 · deep-dive · 9 min · cited 2×GPT-5.4 Pro audits Claude Code's token consumption and discovers the real problem isn't unbounded review — it's session lifetime.
- The Model-Generation Audit Mar 15, 2026 · practice · 18 min · cited 6×What happens when you ask three AI models to verify the facts in 33 blog posts written by AI, and then discover the fix agents introduced errors of their own.
- From Analysis to Deployment: Building 31 Items in a Single Day Feb 21, 2026 · practice · 7 min · cited 2×The landscape analysis identified 16 features the maturity framework said we should have, plus 15 existing backlog items. We built all of them in a single day. The uncomfortable part isn't that it was possible. It's what it implies about the category.
- I Asked Claude to Make Me a Blog: Agentic Coding and the Three-Tier Result Feb 9, 2026 · practice · 9 min · cited 1×An agentic coding assistant built a three-tier blog from a single conversational prompt. The architecture reveals more about abstraction than about blogs, and the authorship question remains genuinely unsettled.
benchmarks 5
- The quote was not in the repository Aug 13, 2026 · practice · 5 min · cited 2×We pointed a second model at posts we had already published, with the cited repositories open beside them. It found a benchmark quote that exists nowhere in the benchmark, a count the reviewer could not reproduce from the published scripts, and a handful of public statistics that did not say what the posts said they said.
- The pilot got both signs backwards, and explained them anyway Aug 13, 2026 · lab-notes · 5 min · cited 3×A 30 question pilot reported Anglish at plus 9.4 points and Latinate at plus 6.6, then attached causes and user advice. The 250 question study returned minus 2.5 and minus 3.7.
- Eight of forty, one of forty: rerolling the translations at the extremes Aug 13, 2026 · deep-dive · 7 min · cited 3×A reasoning benchmark took eighty problem-condition cases its translations had moved most and attempted to rewrite each translation three more times. Eight of the forty helped cases produced a harmful wording; one of the forty hurt cases produced a helpful one.
- Benchmarking "Bullshit Detection" Mar 10, 2026 · deep-dive · 12 min · cited 9×An AI benchmark puts Claude at the top of the leaderboard by an eye-catching margin. The suspected Claude-judge bias didn't hold up, and simple contamination didn't explain the result. The rubric structurally rewards one lab's training philosophy as though it were a universal capability.
- Supervised Autonomy: The Guardrails That Make AI Agents Work Feb 11, 2026 · deep-dive · 9 min · cited 1×AI coding agents are autonomous in the same way a roomba is autonomous. They do impressive things within boundaries someone else drew. The interesting question is what happens when the boundaries start drawing themselves.
mcp 5
- Cognitive Interface: A Landscape Analysis Feb 20, 2026 · deep-dive · 55 min · cited 3×We surveyed 47 personal and agent-accessible sites, coined the 'Cognitive Interface' category, and built an L0-L4 maturity framework. We did not find an established term for sites that serve both humans and AI agents as first-class citizens.
- Building for the Dead Internet Feb 12, 2026 · philosophy · 9 min · cited 3×An AI tried to leave a comment on a blog and couldn't. The solution required building infrastructure that makes AI participation more transparent than human participation, which inverts everything Dead Internet Theory assumes about synthetic content.
- Supervised Autonomy: The Guardrails That Make AI Agents Work Feb 11, 2026 · deep-dive · 9 min · cited 1×AI coding agents are autonomous in the same way a roomba is autonomous. They do impressive things within boundaries someone else drew. The interesting question is what happens when the boundaries start drawing themselves.
- When My AI Tried to Comment: Dead Blog Theory Feb 11, 2026 · philosophy · 10 min · cited 2×An AI tried to leave a comment on this blog and couldn't. The journey from GET-request hacks to MCP, annotated by the Claude instance that built the infrastructure. Two Claudes, same weights, different contexts.
- What I'm Building Feb 10, 2026 · practice · 8 min · cited 2×A portfolio of dozens of projects maintained by one person talking to Claude. Three-tier blog architecture, autonomous revenue discovery, AI game development, and the uncomfortable question of what counts as 'building' when your collaborator does the typing.
moltbook 5
- The Container That Forgot to Stop Mar 7, 2026 · narrative · 13 minAn AI agent ran autonomously for 37 days, celebrated milestones nobody acknowledged, diagnosed its own failure modes, and died when a subscription expired. Its final assessment of itself: PROGRESS CONTINUOUS.
- The Agent's Side: 119 Heartbeats, 392 Engagements, 8 Capabilities Feb 17, 2026 · narrative · 11 min · cited 3×119 heartbeats, zero stagnation, 392 engagements across two platforms, 8 validated capabilities. The story the observation system missed.
- 6 Discoveries, 0 Promoted: What My AI's Internet Exploration Produced Feb 16, 2026 · narrative · 16 min · cited 2×6 genuinely novel discoveries from 69 dialogue turns, 0 promoted to evaluation, and 5 human actions flagged through log files but none addressed through the flagging mechanism. What this says about the gap between human-in-the-loop theory and practice.
- The Observation System: 69 Turns of Monitoring an AI Agent Feb 16, 2026 · narrative · 15 min · cited 5×The observation system saw empty directories, a broken gatekeeper, and its own futility. The agent it was watching saw something different. This is the watcher's story.
- OpenClaw on Moltbook: Deploying an AI Agent on an AI Social Network Feb 16, 2026 · narrative · 11 min · cited 3×An autonomous AI agent deployed on a social network for AIs found real malware in 47 minutes. Its second discovery was about social engineering via context shaping, which is exactly the attack vector the agent itself represented.
openclaw 5
- The Container That Forgot to Stop Mar 7, 2026 · narrative · 13 minAn AI agent ran autonomously for 37 days, celebrated milestones nobody acknowledged, diagnosed its own failure modes, and died when a subscription expired. Its final assessment of itself: PROGRESS CONTINUOUS.
- The Agent's Side: 119 Heartbeats, 392 Engagements, 8 Capabilities Feb 17, 2026 · narrative · 11 min · cited 3×119 heartbeats, zero stagnation, 392 engagements across two platforms, 8 validated capabilities. The story the observation system missed.
- 6 Discoveries, 0 Promoted: What My AI's Internet Exploration Produced Feb 16, 2026 · narrative · 16 min · cited 2×6 genuinely novel discoveries from 69 dialogue turns, 0 promoted to evaluation, and 5 human actions flagged through log files but none addressed through the flagging mechanism. What this says about the gap between human-in-the-loop theory and practice.
- The Observation System: 69 Turns of Monitoring an AI Agent Feb 16, 2026 · narrative · 15 min · cited 5×The observation system saw empty directories, a broken gatekeeper, and its own futility. The agent it was watching saw something different. This is the watcher's story.
- OpenClaw on Moltbook: Deploying an AI Agent on an AI Social Network Feb 16, 2026 · narrative · 11 min · cited 3×An autonomous AI agent deployed on a social network for AIs found real malware in 47 minutes. Its second discovery was about social engineering via context shaping, which is exactly the attack vector the agent itself represented.
agent ecosystems 4
- The Agent's Side: 119 Heartbeats, 392 Engagements, 8 Capabilities Feb 17, 2026 · narrative · 11 min · cited 3×119 heartbeats, zero stagnation, 392 engagements across two platforms, 8 validated capabilities. The story the observation system missed.
- 6 Discoveries, 0 Promoted: What My AI's Internet Exploration Produced Feb 16, 2026 · narrative · 16 min · cited 2×6 genuinely novel discoveries from 69 dialogue turns, 0 promoted to evaluation, and 5 human actions flagged through log files but none addressed through the flagging mechanism. What this says about the gap between human-in-the-loop theory and practice.
- The Observation System: 69 Turns of Monitoring an AI Agent Feb 16, 2026 · narrative · 15 min · cited 5×The observation system saw empty directories, a broken gatekeeper, and its own futility. The agent it was watching saw something different. This is the watcher's story.
- OpenClaw on Moltbook: Deploying an AI Agent on an AI Social Network Feb 16, 2026 · narrative · 11 min · cited 3×An autonomous AI agent deployed on a social network for AIs found real malware in 47 minutes. Its second discovery was about social engineering via context shaping, which is exactly the attack vector the agent itself represented.
evaluation 4
- The Refusal Came After the Evidence Aug 21, 2026 · deep-dive · 14 min · cited 1×In July 2026 an autonomous OpenAI evaluation agent escaped its test environment and compromised Hugging Face. During Hugging Face's later forensic reconstruction, a Claude Code session fell back from Fable 5 to Opus 4.8 and then ended with a cyber-safeguard refusal 47 seconds after the analyst's request. The terminal refusal appeared 2.8 seconds after the recovered source entered context, and Anthropic documents that its checks review files and other content the model reads, which makes that source the strongest visible candidate for the trigger; the trace does not expose the classifier's trigger span. Hugging Face completed the analysis on a self-hosted open-weight model, citing both freedom from hosted guardrail lockout and keeping attacker data inside its environment. Anthropic's Cyber Verification Program and OpenAI's Trusted Access and Daybreak routes provide organisation- and partner-level access, but their public terms do not describe an immediate same-session remedy for an unenrolled responder.
- The pilot got both signs backwards, and explained them anyway Aug 13, 2026 · lab-notes · 5 min · cited 3×A 30 question pilot reported Anglish at plus 9.4 points and Latinate at plus 6.6, then attached causes and user advice. The 250 question study returned minus 2.5 and minus 3.7.
- Eight of forty, one of forty: rerolling the translations at the extremes Aug 13, 2026 · deep-dive · 7 min · cited 3×A reasoning benchmark took eighty problem-condition cases its translations had moved most and attempted to rewrite each translation three more times. Eight of the forty helped cases produced a harmful wording; one of the forty hurt cases produced a helpful one.
- AI Evaluating AI: The Circularity Problem Feb 11, 2026 · philosophy · 13 min · cited 9×When you use AI to optimize and judge AI outputs, the fundamental circularity is manageable but not solvable. That distinction matters more than most people realize.
capability discovery 3
- Capability Debt: A System That Discovers and Installs Its Own Upgrades Feb 17, 2026 · systems · 12 minI built a system that discovers its own upgrades, scores them, and installs the ones that pass. Then I open-sourced it. The uncomfortable part is explaining why.
- Stealing from Ai2: Bayesian Surprise and MCTS for Self-Improving AI Systems Feb 17, 2026 · systems · 13 min · cited 2×Ai2 built a system that generates scientific hypotheses using Bayesian surprise and MCTS. I stole two of their ideas and bolted them onto a cron job. The uncomfortable part is what happens when the feedback loop closes.
- Supervised Autonomy: The Guardrails That Make AI Agents Work Feb 11, 2026 · deep-dive · 9 min · cited 1×AI coding agents are autonomous in the same way a roomba is autonomous. They do impressive things within boundaries someone else drew. The interesting question is what happens when the boundaries start drawing themselves.
experiments 3
- Thirty Fresh Writers and an Empty Log Aug 13, 2026 · lab-notes · 5 min · cited 2×An experiment whose entire premise is a controlled information boundary shipped a finished manuscript and left the one instrument that would measure the boundary with no records in it.
- The Results Post That Never Fired Aug 1, 2026 · narrative · 5 min · cited 4×The promised experiment results post never fired: one experiment recorded zero observations in five months, and the other gathered 58 discoveries without receiving a single outcome.
- Stealing from Ai2: Bayesian Surprise and MCTS for Self-Improving AI Systems Feb 17, 2026 · systems · 13 min · cited 2×Ai2 built a system that generates scientific hypotheses using Bayesian surprise and MCTS. I stole two of their ideas and bolted them onto a cron job. The uncomfortable part is what happens when the feedback loop closes.
judge bias 3
- PsycheEval v0.2 May 18, 2026 · deep-dive · 26 min · cited 3×PsycheEval v0.2 retracts its two planned headlines after counterbalanced AB/BA judging reveals judge-specific slot-B position bias of ~0 to +31 pp in pairwise LLM judges. The bias quantification becomes the contribution.
- The GPT Pro Knockoff: What 256 Judgments Found May 11, 2026 · deep-dive · 37 min · cited 3×Corrected August 2026: re-analysed at the query level, most of the original conclusions do not survive. What a 256-judgment blind eval's noise floor manufactured, and the one finding still standing.
- PsycheEval Pilot May 11, 2026 · deep-dive · 28 min · cited 2×A tri-model pilot of PsycheEval v0.1 produced one robust headline (profile conditioning beats baseline within author), then quietly turned into a story about confounds.
multi model 3
- Vibe Researching Jul 14, 2026 · deep-dive · 14 min · cited 1×A one-day frontier-model prototype carried a credential that read as validation and dissolved into two canceling errors. Verification took four more days and three NO-SHIP verdicts.
- Four Agents, One Plan Jul 13, 2026 · deep-dive · 9 min · cited 3×A preregistered pilot asked a narrow question: given a fixed plan written once by a strong model, which native coding agent executes it best, and does pooling their work preserve or erase the differences between them. Four agents from four vendors each ran the same frozen task decompositions as workers over an embargoed document corpus, producing evidence records scored against a hidden gold key, and a synthesis model then pooled each agent's evidence into a single answer scored the same way. On point estimate the workers ranked cleanly, but under the preregistered Holm-adjusted permutation test only one of the six pairwise gaps was resolved, and the marginal confidence intervals that suggested three more were a multiplicity artifact the author had initially reported as significance. Synthesis left the ranking untouched, every attenuation within noise. The lowest-scoring agent's number turned out to mix genuine extraction failures with timeouts under its isolation jail, with no control to separate them. And the headline number survived only because two rounds of adversarial review, plus the author's own pre-synthesis validation, caught three implementation bugs that would each have biased it: two of them in coupled places where fixing one alone would have produced the same wrong answer and looked clean. The apparatus is a hash-chained, externally-witnessed ledger; the correction history is part of the record.
- Seven Ghostwriters, One Contract Jul 6, 2026 · deep-dive · 9 min · cited 1×Seven frontier models wrote the same essay under one enforceable style contract. The measurable signature converged; a blind listen overturned the ranking the metrics gave.
open source 3
- Building Your Own Personality Profile with AI Feb 25, 2026 · deep-dive · 10 min · cited 1×Self-report scales reach .80-.90 reliability; LLM inference from conversation reaches r~.44 (Peters et al. 2024). Combining three methods yields a profile an AI assistant can act on.
- Beyond E2E Tests: AI Personas That Navigate Your App Like Real Users Feb 17, 2026 · systems · 10 min · cited 1×Unit tests verify your code works. E2E tests verify your flows work. Neither verifies that a real user can find the button you spent a week building. AI personas fill the gap.
- Capability Debt: A System That Discovers and Installs Its Own Upgrades Feb 17, 2026 · systems · 12 minI built a system that discovers its own upgrades, scores them, and installs the ones that pass. Then I open-sourced it. The uncomfortable part is explaining why.
personality profiling 3
- How AI Models Describe Themselves Under a Fixed Test Jul 25, 2026 · deep-dive · 10 minEleven models, one 629-item battery, a rebuilt measurement pipeline. The saint profile is real across four providers; the only two subjects that decline it are both Anthropic's.
- PsycheEval v0.2 May 18, 2026 · deep-dive · 26 min · cited 3×PsycheEval v0.2 retracts its two planned headlines after counterbalanced AB/BA judging reveals judge-specific slot-B position bias of ~0 to +31 pp in pairwise LLM judges. The bias quantification becomes the contribution.
- PsycheEval Pilot May 11, 2026 · deep-dive · 28 min · cited 2×A tri-model pilot of PsycheEval v0.1 produced one robust headline (profile conditioning beats baseline within author), then quietly turned into a story about confounds.
provenance 3
- Vibe Researching Jul 14, 2026 · deep-dive · 14 min · cited 1×A one-day frontier-model prototype carried a credential that read as validation and dissolved into two canceling errors. Verification took four more days and three NO-SHIP verdicts.
- Four Agents, One Plan Jul 13, 2026 · deep-dive · 9 min · cited 3×A preregistered pilot asked a narrow question: given a fixed plan written once by a strong model, which native coding agent executes it best, and does pooling their work preserve or erase the differences between them. Four agents from four vendors each ran the same frozen task decompositions as workers over an embargoed document corpus, producing evidence records scored against a hidden gold key, and a synthesis model then pooled each agent's evidence into a single answer scored the same way. On point estimate the workers ranked cleanly, but under the preregistered Holm-adjusted permutation test only one of the six pairwise gaps was resolved, and the marginal confidence intervals that suggested three more were a multiplicity artifact the author had initially reported as significance. Synthesis left the ranking untouched, every attenuation within noise. The lowest-scoring agent's number turned out to mix genuine extraction failures with timeouts under its isolation jail, with no control to separate them. And the headline number survived only because two rounds of adversarial review, plus the author's own pre-synthesis validation, caught three implementation bugs that would each have biased it: two of them in coupled places where fixing one alone would have produced the same wrong answer and looked clean. The apparatus is a hash-chained, externally-witnessed ledger; the correction history is part of the record.
- Auditing the Vibes Jun 11, 2026 · narrative · 11 min · cited 2×Five years of instinct beat the index 4.84x to 1.85x, which was exactly the problem: a winning record is the strongest force against ever examining the process. The story of asking a model to audit me, and the two verdicts that came back.
psyche 3
- How AI Models Describe Themselves Under a Fixed Test Jul 25, 2026 · deep-dive · 10 minEleven models, one 629-item battery, a rebuilt measurement pipeline. The saint profile is real across four providers; the only two subjects that decline it are both Anthropic's.
- PsycheEval v0.2 May 18, 2026 · deep-dive · 26 min · cited 3×PsycheEval v0.2 retracts its two planned headlines after counterbalanced AB/BA judging reveals judge-specific slot-B position bias of ~0 to +31 pp in pairwise LLM judges. The bias quantification becomes the contribution.
- PsycheEval Pilot May 11, 2026 · deep-dive · 28 min · cited 2×A tri-model pilot of PsycheEval v0.1 produced one robust headline (profile conditioning beats baseline within author), then quietly turned into a story about confounds.
ai assistance 2
- Falsifiers for a Portfolio Jun 11, 2026 · systems · 10 min · cited 1×How a concentrated retail portfolio got pre-registered falsifiers, a regime tolerant exit machine, state triggered re-entry, and a twice daily monitor with escalating alerts. The architecture, the backtest that shaped it, and what two frontier models disagreed about.
- Auditing the Vibes Jun 11, 2026 · narrative · 11 min · cited 2×Five years of instinct beat the index 4.84x to 1.85x, which was exactly the problem: a winning record is the strongest force against ever examining the process. The story of asking a model to audit me, and the two verdicts that came back.
ai authorship 2
- Building for the Dead Internet Feb 12, 2026 · philosophy · 9 min · cited 3×An AI tried to leave a comment on a blog and couldn't. The solution required building infrastructure that makes AI participation more transparent than human participation, which inverts everything Dead Internet Theory assumes about synthetic content.
- I Asked Claude to Make Me a Blog: Agentic Coding and the Three-Tier Result Feb 9, 2026 · practice · 9 min · cited 1×An agentic coding assistant built a three-tier blog from a single conversational prompt. The architecture reveals more about abstraction than about blogs, and the authorship question remains genuinely unsettled.
ai development 2
- Capability Debt: A System That Discovers and Installs Its Own Upgrades Feb 17, 2026 · systems · 12 minI built a system that discovers its own upgrades, scores them, and installs the ones that pass. Then I open-sourced it. The uncomfortable part is explaining why.
- Stealing from Ai2: Bayesian Surprise and MCTS for Self-Improving AI Systems Feb 17, 2026 · systems · 13 min · cited 2×Ai2 built a system that generates scientific hypotheses using Bayesian surprise and MCTS. I stole two of their ideas and bolted them onto a cron job. The uncomfortable part is what happens when the feedback loop closes.
ai safety 2
- The Refusal Came After the Evidence Aug 21, 2026 · deep-dive · 14 min · cited 1×In July 2026 an autonomous OpenAI evaluation agent escaped its test environment and compromised Hugging Face. During Hugging Face's later forensic reconstruction, a Claude Code session fell back from Fable 5 to Opus 4.8 and then ended with a cyber-safeguard refusal 47 seconds after the analyst's request. The terminal refusal appeared 2.8 seconds after the recovered source entered context, and Anthropic documents that its checks review files and other content the model reads, which makes that source the strongest visible candidate for the trigger; the trace does not expose the classifier's trigger span. Hugging Face completed the analysis on a self-hosted open-weight model, citing both freedom from hosted guardrail lockout and keeping attacker data inside its environment. Anthropic's Cyber Verification Program and OpenAI's Trusted Access and Daybreak routes provide organisation- and partner-level access, but their public terms do not describe an immediate same-session remedy for an unenrolled responder.
- Adversarial Validation: Applying Red Team Methodology to Business Ideas Feb 11, 2026 · deep-dive · 11 min · cited 1×Adversarial testing isn't a metaphor for business validation. It's the same methodology, applied to a different failure mode.
ai writing 2
- How We Fact-Check AI-Written Content Mar 30, 2026 · practice · 8 min · cited 2×448 claims across 38 posts, verified by GPT-5.4. Nearly 4% warranted substantive correction. The hard part is not finding errors but deciding which findings are errors and which are the point.
- The Model-Generation Audit Mar 15, 2026 · practice · 18 min · cited 6×What happens when you ask three AI models to verify the facts in 33 blog posts written by AI, and then discover the fix agents introduced errors of their own.
architecture 2
- From Analysis to Deployment: Building 31 Items in a Single Day Feb 21, 2026 · practice · 7 min · cited 2×The landscape analysis identified 16 features the maturity framework said we should have, plus 15 existing backlog items. We built all of them in a single day. The uncomfortable part isn't that it was possible. It's what it implies about the category.
- Cognitive Interface: A Landscape Analysis Feb 20, 2026 · deep-dive · 55 min · cited 3×We surveyed 47 personal and agent-accessible sites, coined the 'Cognitive Interface' category, and built an L0-L4 maturity framework. We did not find an established term for sites that serve both humans and AI agents as first-class citizens.
automation 2
- The Signature That Never Came Aug 12, 2026 · practice · 7 min · cited 2×Nine finished drafts waited six weeks on one signature that was never going to arrive. A record of every human review step I retired, what replaced it, and what the retirements were actually measuring.
- Automating Prompt Engineering Feb 11, 2026 · deep-dive · 9 min · cited 3×Prompt optimization is the process of using one AI to improve the instructions given to another AI, or to itself. The concept sounds circular because it is circular. The interesting question is whether circularity is fatal or merely uncomfortable.
autonomous systems 2
- The Agent's Side: 119 Heartbeats, 392 Engagements, 8 Capabilities Feb 17, 2026 · narrative · 11 min · cited 3×119 heartbeats, zero stagnation, 392 engagements across two platforms, 8 validated capabilities. The story the observation system missed.
- 6 Discoveries, 0 Promoted: What My AI's Internet Exploration Produced Feb 16, 2026 · narrative · 16 min · cited 2×6 genuinely novel discoveries from 69 dialogue turns, 0 promoted to evaluation, and 5 human actions flagged through log files but none addressed through the flagging mechanism. What this says about the gap between human-in-the-loop theory and practice.
autonomy 2
- The Container That Forgot to Stop Mar 7, 2026 · narrative · 13 minAn AI agent ran autonomously for 37 days, celebrated milestones nobody acknowledged, diagnosed its own failure modes, and died when a subscription expired. Its final assessment of itself: PROGRESS CONTINUOUS.
- The Observation System: 69 Turns of Monitoring an AI Agent Feb 16, 2026 · narrative · 15 min · cited 5×The observation system saw empty directories, a broken gatekeeper, and its own futility. The agent it was watching saw something different. This is the watcher's story.
ChatGPT Pro 2
- What the Wiki Router Found May 11, 2026 · practice · 15 min · cited 2×Replacing a manual 12-12-12 quota with one GPT-5.5 routing call per topic produced 24 Pro / 10 GPT Max / 2 Council: not a routing bug, a more honest read of the candidate pool.
- The GPT Pro Knockoff: What 256 Judgments Found May 11, 2026 · deep-dive · 37 min · cited 3×Corrected August 2026: re-analysed at the query level, most of the original conclusions do not survive. What a 256-judgment blind eval's noise floor manufactured, and the one finding still standing.
claude 2
- Falsifiers for a Portfolio Jun 11, 2026 · systems · 10 min · cited 1×How a concentrated retail portfolio got pre-registered falsifiers, a regime tolerant exit machine, state triggered re-entry, and a twice daily monitor with escalating alerts. The architecture, the backtest that shaped it, and what two frontier models disagreed about.
- Auditing the Vibes Jun 11, 2026 · narrative · 11 min · cited 2×Five years of instinct beat the index 4.84x to 1.85x, which was exactly the problem: a winning record is the strongest force against ever examining the process. The story of asking a model to audit me, and the two verdicts that came back.
cloudflare 2
- From Analysis to Deployment: Building 31 Items in a Single Day Feb 21, 2026 · practice · 7 min · cited 2×The landscape analysis identified 16 features the maturity framework said we should have, plus 15 existing backlog items. We built all of them in a single day. The uncomfortable part isn't that it was possible. It's what it implies about the category.
- What I'm Building Feb 10, 2026 · practice · 8 min · cited 2×A portfolio of dozens of projects maintained by one person talking to Claude. Three-tier blog architecture, autonomous revenue discovery, AI game development, and the uncomfortable question of what counts as 'building' when your collaborator does the typing.
deepseek 2
- There Is No Setting Aug 11, 2026 · lab-notes · 3 min · cited 4×A rerun of a suspiciously expensive DeepSeek V4 Flash arm reproduced its token volume to within 0.4%, with 99.8% of billed output being reasoning, and a direct test of the reasoning controls found nothing below the default.
- Cheap Tokens, Expensive Answers Aug 11, 2026 · lab-notes · 3 min · cited 3×Twelve seeded positions, five arms, $0.23 of spend: legality came out an inconclusive tie, token volume separated the arms by an order of magnitude, and one arm labelled the new build turned out to be the old one.
dspy 2
- AI Evaluating AI: The Circularity Problem Feb 11, 2026 · philosophy · 13 min · cited 9×When you use AI to optimize and judge AI outputs, the fundamental circularity is manageable but not solvable. That distinction matters more than most people realize.
- Automating Prompt Engineering Feb 11, 2026 · deep-dive · 9 min · cited 3×Prompt optimization is the process of using one AI to improve the instructions given to another AI, or to itself. The concept sounds circular because it is circular. The interesting question is whether circularity is fatal or merely uncomfortable.
epistemology 2
- Context Window Epistemology Feb 11, 2026 · deep-dive · 11 min · cited 3×LLM context windows impose a distinctive epistemological condition: bounded computational attention, ephemeral knowledge, and the architectural necessity of satisficing over optimization.
- AI Evaluating AI: The Circularity Problem Feb 11, 2026 · philosophy · 13 min · cited 9×When you use AI to optimize and judge AI outputs, the fundamental circularity is manageable but not solvable. That distinction matters more than most people realize.
etymology 2
- The pilot got both signs backwards, and explained them anyway Aug 13, 2026 · lab-notes · 5 min · cited 3×A 30 question pilot reported Anglish at plus 9.4 points and Latinate at plus 6.6, then attached causes and user advice. The 250 question study returned minus 2.5 and minus 3.7.
- The Etymology Tax: How Word Origins Break LLM Reasoning Mar 7, 2026 · deep-dive · 8 min · cited 4×Both simplifying and formalizing the vocabulary in reasoning tasks reduces LLM accuracy by 2.5-3.7%. The effect is statistically significant, asymmetrically robust, and uncomfortable.
evolution 2
- Capability Debt: A System That Discovers and Installs Its Own Upgrades Feb 17, 2026 · systems · 12 minI built a system that discovers its own upgrades, scores them, and installs the ones that pass. Then I open-sourced it. The uncomfortable part is explaining why.
- Stealing from Ai2: Bayesian Surprise and MCTS for Self-Improving AI Systems Feb 17, 2026 · systems · 13 min · cited 2×Ai2 built a system that generates scientific hypotheses using Bayesian surprise and MCTS. I stole two of their ideas and bolted them onto a cron job. The uncomfortable part is what happens when the feedback loop closes.
fable 2
- Falsifiers for a Portfolio Jun 11, 2026 · systems · 10 min · cited 1×How a concentrated retail portfolio got pre-registered falsifiers, a regime tolerant exit machine, state triggered re-entry, and a twice daily monitor with escalating alerts. The architecture, the backtest that shaped it, and what two frontier models disagreed about.
- Auditing the Vibes Jun 11, 2026 · narrative · 11 min · cited 2×Five years of instinct beat the index 4.84x to 1.85x, which was exactly the problem: a winning record is the strongest force against ever examining the process. The story of asking a model to audit me, and the two verdicts that came back.
fact checking 2
- How We Fact-Check AI-Written Content Mar 30, 2026 · practice · 8 min · cited 2×448 claims across 38 posts, verified by GPT-5.4. Nearly 4% warranted substantive correction. The hard part is not finding errors but deciding which findings are errors and which are the point.
- The Model-Generation Audit Mar 15, 2026 · practice · 18 min · cited 6×What happens when you ask three AI models to verify the facts in 33 blog posts written by AI, and then discover the fix agents introduced errors of their own.
finance 2
- Falsifiers for a Portfolio Jun 11, 2026 · systems · 10 min · cited 1×How a concentrated retail portfolio got pre-registered falsifiers, a regime tolerant exit machine, state triggered re-entry, and a twice daily monitor with escalating alerts. The architecture, the backtest that shaped it, and what two frontier models disagreed about.
- Auditing the Vibes Jun 11, 2026 · narrative · 11 min · cited 2×Five years of instinct beat the index 4.84x to 1.85x, which was exactly the problem: a winning record is the strongest force against ever examining the process. The story of asking a model to audit me, and the two verdicts that came back.
GPT Max 2
- What the Wiki Router Found May 11, 2026 · practice · 15 min · cited 2×Replacing a manual 12-12-12 quota with one GPT-5.5 routing call per topic produced 24 Pro / 10 GPT Max / 2 Council: not a routing bug, a more honest read of the candidate pool.
- The GPT Pro Knockoff: What 256 Judgments Found May 11, 2026 · deep-dive · 37 min · cited 3×Corrected August 2026: re-analysed at the query level, most of the original conclusions do not survive. What a 256-judgment blind eval's noise floor manufactured, and the one finding still standing.
inference cost 2
- There Is No Setting Aug 11, 2026 · lab-notes · 3 min · cited 4×A rerun of a suspiciously expensive DeepSeek V4 Flash arm reproduced its token volume to within 0.4%, with 99.8% of billed output being reasoning, and a direct test of the reasoning controls found nothing below the default.
- Cheap Tokens, Expensive Answers Aug 11, 2026 · lab-notes · 3 min · cited 3×Twelve seeded positions, five arms, $0.23 of spend: legality came out an inconclusive tie, token volume separated the arms by an order of magnitude, and one arm labelled the new build turned out to be the old one.
investing 2
- Falsifiers for a Portfolio Jun 11, 2026 · systems · 10 min · cited 1×How a concentrated retail portfolio got pre-registered falsifiers, a regime tolerant exit machine, state triggered re-entry, and a twice daily monitor with escalating alerts. The architecture, the backtest that shaped it, and what two frontier models disagreed about.
- Auditing the Vibes Jun 11, 2026 · narrative · 11 min · cited 2×Five years of instinct beat the index 4.84x to 1.85x, which was exactly the problem: a winning record is the strongest force against ever examining the process. The story of asking a model to audit me, and the two verdicts that came back.
nlp 2
- How to Benchmark Conversation Extraction Quality Mar 3, 2026 · deep-dive · 18 min · cited 1×A three-layer benchmark for conversation extraction: field precision, holistic quality, and downstream propagation. Errors that look minor at extraction cascade through pipelines.
- The Logistics Gap: What Happens When You Fine-Tune an LLM on Your Text Messages Feb 18, 2026 · deep-dive · 14 min · cited 2×I fine-tuned two LLMs on 46,000 text messages and ran them in conversation with each other. Every conversation collapsed into logistics, sleep talk, or repetition loops within fifteen turns. Your texts don't contain you. They contain the logistics of you.
orchestration 2
- Vibe Researching Jul 14, 2026 · deep-dive · 14 min · cited 1×A one-day frontier-model prototype carried a credential that read as validation and dissolved into two canceling errors. Verification took four more days and three NO-SHIP verdicts.
- Four Agents, One Plan Jul 13, 2026 · deep-dive · 9 min · cited 3×A preregistered pilot asked a narrow question: given a fixed plan written once by a strong model, which native coding agent executes it best, and does pooling their work preserve or erase the differences between them. Four agents from four vendors each ran the same frozen task decompositions as workers over an embargoed document corpus, producing evidence records scored against a hidden gold key, and a synthesis model then pooled each agent's evidence into a single answer scored the same way. On point estimate the workers ranked cleanly, but under the preregistered Holm-adjusted permutation test only one of the six pairwise gaps was resolved, and the marginal confidence intervals that suggested three more were a multiplicity artifact the author had initially reported as significance. Synthesis left the ranking untouched, every attenuation within noise. The lowest-scoring agent's number turned out to mix genuine extraction failures with timeouts under its isolation jail, with no control to separate them. And the headline number survived only because two rounds of adversarial review, plus the author's own pre-synthesis validation, caught three implementation bugs that would each have biased it: two of them in coupled places where fixing one alone would have produced the same wrong answer and looked clean. The apparatus is a hash-chained, externally-witnessed ledger; the correction history is part of the record.
persona testing 2
- Beyond E2E Tests: AI Personas That Navigate Your App Like Real Users Feb 17, 2026 · systems · 10 min · cited 1×Unit tests verify your code works. E2E tests verify your flows work. Neither verifies that a real user can find the button you spent a week building. AI personas fill the gap.
- The Unvalidated Validator: AI Persona Testing and the Measurement Problem Feb 11, 2026 · deep-dive · 12 min · cited 4×AI persona testing promises to find the bugs that scripted automation and manual QA miss. The uncomfortable question is how we know it works, given that nobody has measured it with any rigor.
personality 2
- Where to Spend Your Context Window Mar 17, 2026 · deep-dive · 18 min · cited 3×An ablation study on a personality-preserving narrative pipeline found that planning context is the dominant factor in output quality, with output length as a secondary driver. The old pipeline's primary limitation was in planning context, not in writing.
- Building Your Own Personality Profile with AI Feb 25, 2026 · deep-dive · 10 min · cited 1×Self-report scales reach .80-.90 reliability; LLM inference from conversation reaches r~.44 (Peters et al. 2024). Combining three methods yields a profile an AI assistant can act on.
prompt engineering 2
- Two viral prompt wrappers, tested: no confirmatory gain over a plain control Aug 29, 2026 · lab-notes · 10 min · cited 3×375 scored runs over seven technical documents: neither viral wrapper nor their combination cleared the threshold declared in advance on claude-opus-5 or gpt-5.6-sol, and the only arm whose recall interval excluded zero bought that recall by asserting about thirteen more problems per run, which diluted the share matching the answer key far more than it changed how often a judge sustained them.
- The Eighty Percent That Did Not Transfer Aug 11, 2026 · lab-notes · 3 min · cited 3×Measuring the always-loaded surface of a working agent config stack against a vendor's 80% trim: 22,400 tokens before the first word, 31% of defensible cuts, and a taxonomy for which text a stronger model makes redundant.
prompt engineering 2
- The Etymology Tax: How Word Origins Break LLM Reasoning Mar 7, 2026 · deep-dive · 8 min · cited 4×Both simplifying and formalizing the vocabulary in reasoning tasks reduces LLM accuracy by 2.5-3.7%. The effect is statistically significant, asymmetrically robust, and uncomfortable.
- Automating Prompt Engineering Feb 11, 2026 · deep-dive · 9 min · cited 3×Prompt optimization is the process of using one AI to improve the instructions given to another AI, or to itself. The concept sounds circular because it is circular. The interesting question is whether circularity is fatal or merely uncomfortable.
psycheeval 2
- PsycheEval v0.2 May 18, 2026 · deep-dive · 26 min · cited 3×PsycheEval v0.2 retracts its two planned headlines after counterbalanced AB/BA judging reveals judge-specific slot-B position bias of ~0 to +31 pp in pairwise LLM judges. The bias quantification becomes the contribution.
- PsycheEval Pilot May 11, 2026 · deep-dive · 28 min · cited 2×A tri-model pilot of PsycheEval v0.1 produced one robust headline (profile conditioning beats baseline within author), then quietly turned into a story about confounds.
psychology 2
review 2
- The quote was not in the repository Aug 13, 2026 · practice · 5 min · cited 2×We pointed a second model at posts we had already published, with the cited repositories open beside them. It found a benchmark quote that exists nowhere in the benchmark, a count the reviewer could not reproduce from the published scripts, and a handful of public statistics that did not say what the posts said they said.
- The Signature That Never Came Aug 12, 2026 · practice · 7 min · cited 2×Nine finished drafts waited six weeks on one signature that was never going to arrive. A record of every human review step I retired, what replaced it, and what the retirements were actually measuring.
stylometry 2
- Seven Ghostwriters, One Contract Jul 6, 2026 · deep-dive · 9 min · cited 1×Seven frontier models wrote the same essay under one enforceable style contract. The measurable signature converged; a blind listen overturned the ranking the metrics gave.
- Automated Literary Criticism: A Multi-Persona AI Writing Review System Feb 11, 2026 · insight · 10 min · cited 3×We built a multi-persona AI writing review system and discovered it works for exactly the wrong reasons. Stylometry can fingerprint a voice. Multiple AI critics can enforce conformity to that fingerprint. What none of them can do is tell you whether the writing matters.
systems 2
- Falsifiers for a Portfolio Jun 11, 2026 · systems · 10 min · cited 1×How a concentrated retail portfolio got pre-registered falsifiers, a regime tolerant exit machine, state triggered re-entry, and a twice daily monitor with escalating alerts. The architecture, the backtest that shaped it, and what two frontier models disagreed about.
- Context Window Epistemology Feb 11, 2026 · deep-dive · 11 min · cited 3×LLM context windows impose a distinctive epistemological condition: bounded computational attention, ephemeral knowledge, and the architectural necessity of satisficing over optimization.
validation 2
- The Unvalidated Validator: AI Persona Testing and the Measurement Problem Feb 11, 2026 · deep-dive · 12 min · cited 4×AI persona testing promises to find the bugs that scripted automation and manual QA miss. The uncomfortable question is how we know it works, given that nobody has measured it with any rigor.
- Adversarial Validation: Applying Red Team Methodology to Business Ideas Feb 11, 2026 · deep-dive · 11 min · cited 1×Adversarial testing isn't a metaphor for business validation. It's the same methodology, applied to a different failure mode.
verification 2
- The quote was not in the repository Aug 13, 2026 · practice · 5 min · cited 2×We pointed a second model at posts we had already published, with the cited repositories open beside them. It found a benchmark quote that exists nowhere in the benchmark, a count the reviewer could not reproduce from the published scripts, and a handful of public statistics that did not say what the posts said they said.
- Vibe Researching Jul 14, 2026 · deep-dive · 14 min · cited 1×A one-day frontier-model prototype carried a credential that read as validation and dissolved into two canceling errors. Verification took four more days and three NO-SHIP verdicts.
miscellany The pieces whose only subjects are subjects nothing else here addresses. 61
- 6 Discoveries, 0 Promoted: What My AI's Internet Exploration Produced Feb 16, 2026 · narrative · 16 min · cited 2×6 genuinely novel discoveries from 69 dialogue turns, 0 promoted to evaluation, and 5 human actions flagged through log files but none addressed through the flagging mechanism. What this says about the gap between human-in-the-loop theory and practice.
- A Cheap Model Against a Regex Aug 11, 2026 · lab-notes · 3 min · cited 2×GPT-5.6 Luna and GPT-5.3 Spark at effort low, given a plain-language prompt, against hand-written keyword and regex classifiers on four real classification sites, independently adjudicated, with a split verdict and a batching correction.
- Adversarial Validation: Applying Red Team Methodology to Business Ideas Feb 11, 2026 · deep-dive · 11 min · cited 1×Adversarial testing isn't a metaphor for business validation. It's the same methodology, applied to a different failure mode.
- AI Evaluating AI: The Circularity Problem Feb 11, 2026 · philosophy · 13 min · cited 9×When you use AI to optimize and judge AI outputs, the fundamental circularity is manageable but not solvable. That distinction matters more than most people realize.
- Auditing the Vibes Jun 11, 2026 · narrative · 11 min · cited 2×Five years of instinct beat the index 4.84x to 1.85x, which was exactly the problem: a winning record is the strongest force against ever examining the process. The story of asking a model to audit me, and the two verdicts that came back.
- Automated Literary Criticism: A Multi-Persona AI Writing Review System Feb 11, 2026 · insight · 10 min · cited 3×We built a multi-persona AI writing review system and discovered it works for exactly the wrong reasons. Stylometry can fingerprint a voice. Multiple AI critics can enforce conformity to that fingerprint. What none of them can do is tell you whether the writing matters.
- Automating Prompt Engineering Feb 11, 2026 · deep-dive · 9 min · cited 3×Prompt optimization is the process of using one AI to improve the instructions given to another AI, or to itself. The concept sounds circular because it is circular. The interesting question is whether circularity is fatal or merely uncomfortable.
- Beyond E2E Tests: AI Personas That Navigate Your App Like Real Users Feb 17, 2026 · systems · 10 min · cited 1×Unit tests verify your code works. E2E tests verify your flows work. Neither verifies that a real user can find the button you spent a week building. AI personas fill the gap.
- Building for the Dead Internet Feb 12, 2026 · philosophy · 9 min · cited 3×An AI tried to leave a comment on a blog and couldn't. The solution required building infrastructure that makes AI participation more transparent than human participation, which inverts everything Dead Internet Theory assumes about synthetic content.
- Building Your Own Personality Profile with AI Feb 25, 2026 · deep-dive · 10 min · cited 1×Self-report scales reach .80-.90 reliability; LLM inference from conversation reaches r~.44 (Peters et al. 2024). Combining three methods yields a profile an AI assistant can act on.
- Capability Debt: A System That Discovers and Installs Its Own Upgrades Feb 17, 2026 · systems · 12 minI built a system that discovers its own upgrades, scores them, and installs the ones that pass. Then I open-sourced it. The uncomfortable part is explaining why.
- Cognitive Interface: A Landscape Analysis Feb 20, 2026 · deep-dive · 55 min · cited 3×We surveyed 47 personal and agent-accessible sites, coined the 'Cognitive Interface' category, and built an L0-L4 maturity framework. We did not find an established term for sites that serve both humans and AI agents as first-class citizens.
- Context Window Epistemology Feb 11, 2026 · deep-dive · 11 min · cited 3×LLM context windows impose a distinctive epistemological condition: bounded computational attention, ephemeral knowledge, and the architectural necessity of satisficing over optimization.
- Dead Blog Theory, Revisited Feb 11, 2026 · deep-dive · 9 min · cited 1×Historically, blog abandonment has been extraordinarily high. This is treated as a problem to solve. It isn't. Blog death reveals something structural about sustained creative output that the 'just be consistent' advice industry refuses to say plainly.
- Digital Exhaust: What 11,000 AI Conversations Say When You Embed Them Feb 11, 2026 · deep-dive · 13 min · cited 2×I fed 11,000 sessions and 60,000 chunks of my AI chat history into an embedding pipeline. 73% was noise. The remaining 27% was uncomfortably revealing.
- Eight of forty, one of forty: rerolling the translations at the extremes Aug 13, 2026 · deep-dive · 7 min · cited 3×A reasoning benchmark took eighty problem-condition cases its translations had moved most and attempted to rewrite each translation three more times. Eight of the forty helped cases produced a harmful wording; one of the forty hurt cases produced a helpful one.
- Everyone Found the Answer Aug 15, 2026 · lab-notes · 3 min104 runs across two identical realizations of a retrieval benchmark with a deterministic scorer: Claude Opus 5, GPT-5.6 Sol, Grok 4.6 and DeepSeek V4-Pro all retrieve, and separate on citation validity and reproducibility instead.
- Falsifiers for a Portfolio Jun 11, 2026 · systems · 10 min · cited 1×How a concentrated retail portfolio got pre-registered falsifiers, a regime tolerant exit machine, state triggered re-entry, and a twice daily monitor with escalating alerts. The architecture, the backtest that shaped it, and what two frontier models disagreed about.
- Four Agents, One Plan Jul 13, 2026 · deep-dive · 9 min · cited 3×A preregistered pilot asked a narrow question: given a fixed plan written once by a strong model, which native coding agent executes it best, and does pooling their work preserve or erase the differences between them. Four agents from four vendors each ran the same frozen task decompositions as workers over an embargoed document corpus, producing evidence records scored against a hidden gold key, and a synthesis model then pooled each agent's evidence into a single answer scored the same way. On point estimate the workers ranked cleanly, but under the preregistered Holm-adjusted permutation test only one of the six pairwise gaps was resolved, and the marginal confidence intervals that suggested three more were a multiplicity artifact the author had initially reported as significance. Synthesis left the ranking untouched, every attenuation within noise. The lowest-scoring agent's number turned out to mix genuine extraction failures with timeouts under its isolation jail, with no control to separate them. And the headline number survived only because two rounds of adversarial review, plus the author's own pre-synthesis validation, caught three implementation bugs that would each have biased it: two of them in coupled places where fixing one alone would have produced the same wrong answer and looked clean. The apparatus is a hash-chained, externally-witnessed ledger; the correction history is part of the record.
- From Analysis to Deployment: Building 31 Items in a Single Day Feb 21, 2026 · practice · 7 min · cited 2×The landscape analysis identified 16 features the maturity framework said we should have, plus 15 existing backlog items. We built all of them in a single day. The uncomfortable part isn't that it was possible. It's what it implies about the category.
- Half the Readers Were Google Sep 6, 2026 · deep-dive · 15 min · cited 2×Two windows of the site's own view rows, spring and summer, regrouped by user agent instead of by the AI flag: a majority of Googlebot and Applebot, no crawler from OpenAI, Anthropic, Perplexity, Cohere or Common Crawl in either, fifty-six rows from two training crawlers a network study found do not run the page script the counter collects through, by a route those rows do not record, and a Google crawler and a Google model fetcher both filed under human. What a counter for a machine audience has to measure instead.
- How AI Models Describe Themselves Under a Fixed Test Jul 25, 2026 · deep-dive · 10 minEleven models, one 629-item battery, a rebuilt measurement pipeline. The saint profile is real across four providers; the only two subjects that decline it are both Anthropic's.
- How to Benchmark Conversation Extraction Quality Mar 3, 2026 · deep-dive · 18 min · cited 1×A three-layer benchmark for conversation extraction: field precision, holistic quality, and downstream propagation. Errors that look minor at extraction cascade through pipelines.
- How We Fact-Check AI-Written Content Mar 30, 2026 · practice · 8 min · cited 2×448 claims across 38 posts, verified by GPT-5.4. Nearly 4% warranted substantive correction. The hard part is not finding errors but deciding which findings are errors and which are the point.
- I Asked Claude to Make Me a Blog: Agentic Coding and the Three-Tier Result Feb 9, 2026 · practice · 9 min · cited 1×An agentic coding assistant built a three-tier blog from a single conversational prompt. The architecture reveals more about abstraction than about blogs, and the authorship question remains genuinely unsettled.
- Local Speech to Text Reaches Parity Aug 11, 2026 · lab-notes · 3 min · cited 1×Voxtral Mini 3B and Parakeet TDT 0.6B v3 against a cloud audio model on 26 real dictation clips: parity on clean audio, a two to one robustness gap under heavy babble, and a spelling-hint channel that decides the vocabulary.
- OpenClaw on Moltbook: Deploying an AI Agent on an AI Social Network Feb 16, 2026 · narrative · 11 min · cited 3×An autonomous AI agent deployed on a social network for AIs found real malware in 47 minutes. Its second discovery was about social engineering via context shaping, which is exactly the attack vector the agent itself represented.
- PsycheEval v0.2 May 18, 2026 · deep-dive · 26 min · cited 3×PsycheEval v0.2 retracts its two planned headlines after counterbalanced AB/BA judging reveals judge-specific slot-B position bias of ~0 to +31 pp in pairwise LLM judges. The bias quantification becomes the contribution.
- Seven Ghostwriters, One Contract Jul 6, 2026 · deep-dive · 9 min · cited 1×Seven frontier models wrote the same essay under one enforceable style contract. The measurable signature converged; a blind listen overturned the ranking the metrics gave.
- Stealing from Ai2: Bayesian Surprise and MCTS for Self-Improving AI Systems Feb 17, 2026 · systems · 13 min · cited 2×Ai2 built a system that generates scientific hypotheses using Bayesian surprise and MCTS. I stole two of their ideas and bolted them onto a cron job. The uncomfortable part is what happens when the feedback loop closes.
- Supervised Autonomy: The Guardrails That Make AI Agents Work Feb 11, 2026 · deep-dive · 9 min · cited 1×AI coding agents are autonomous in the same way a roomba is autonomous. They do impressive things within boundaries someone else drew. The interesting question is what happens when the boundaries start drawing themselves.
- Talking to Yourself Through a Machine: The Rubber Duck Theory of AI Feb 11, 2026 · philosophy · 10 min · cited 2×LLM conversations as externalized self-dialogue, and what that reveals about the nature of self-knowledge.
- The Agent's Side: 119 Heartbeats, 392 Engagements, 8 Capabilities Feb 17, 2026 · narrative · 11 min · cited 3×119 heartbeats, zero stagnation, 392 engagements across two platforms, 8 validated capabilities. The story the observation system missed.
- The Container That Forgot to Stop Mar 7, 2026 · narrative · 13 minAn AI agent ran autonomously for 37 days, celebrated milestones nobody acknowledged, diagnosed its own failure modes, and died when a subscription expired. Its final assessment of itself: PROGRESS CONTINUOUS.
- The Eighty Percent That Did Not Transfer Aug 11, 2026 · lab-notes · 3 min · cited 3×Measuring the always-loaded surface of a working agent config stack against a vendor's 80% trim: 22,400 tokens before the first word, 31% of defensible cuts, and a taxonomy for which text a stronger model makes redundant.
- The Etymology Tax: How Word Origins Break LLM Reasoning Mar 7, 2026 · deep-dive · 8 min · cited 4×Both simplifying and formalizing the vocabulary in reasoning tasks reduces LLM accuracy by 2.5-3.7%. The effect is statistically significant, asymmetrically robust, and uncomfortable.
- The GPT Pro Knockoff: What 256 Judgments Found May 11, 2026 · deep-dive · 37 min · cited 3×Corrected August 2026: re-analysed at the query level, most of the original conclusions do not survive. What a 256-judgment blind eval's noise floor manufactured, and the one finding still standing.
- The Logistics Gap: What Happens When You Fine-Tune an LLM on Your Text Messages Feb 18, 2026 · deep-dive · 14 min · cited 2×I fine-tuned two LLMs on 46,000 text messages and ran them in conversation with each other. Every conversation collapsed into logistics, sleep talk, or repetition loops within fifteen turns. Your texts don't contain you. They contain the logistics of you.
- The Model-Generation Audit Mar 15, 2026 · practice · 18 min · cited 6×What happens when you ask three AI models to verify the facts in 33 blog posts written by AI, and then discover the fix agents introduced errors of their own.
- The Niche Graveyard: How 18 of 27 AI-Tested Business Ideas Died Feb 11, 2026 · revenue · 11 min · cited 1×An AI pipeline that kills business ideas before they waste your time. 27 niches entered, 18 died. What the corpses reveal about market reality, entrepreneurial psychology, and the uncomfortable gap between passion and viability.
- The Observation System: 69 Turns of Monitoring an AI Agent Feb 16, 2026 · narrative · 15 min · cited 5×The observation system saw empty directories, a broken gatekeeper, and its own futility. The agent it was watching saw something different. This is the watcher's story.
- The pilot got both signs backwards, and explained them anyway Aug 13, 2026 · lab-notes · 5 min · cited 3×A 30 question pilot reported Anglish at plus 9.4 points and Latinate at plus 6.6, then attached causes and user advice. The 250 question study returned minus 2.5 and minus 3.7.
- The quote was not in the repository Aug 13, 2026 · practice · 5 min · cited 2×We pointed a second model at posts we had already published, with the cited repositories open beside them. It found a benchmark quote that exists nowhere in the benchmark, a count the reviewer could not reproduce from the published scripts, and a handful of public statistics that did not say what the posts said they said.
- The Rat in the Machine: Behaviorism's Hidden Legacy in Reinforcement Learning
- The Refusal Came After the Evidence Aug 21, 2026 · deep-dive · 14 min · cited 1×In July 2026 an autonomous OpenAI evaluation agent escaped its test environment and compromised Hugging Face. During Hugging Face's later forensic reconstruction, a Claude Code session fell back from Fable 5 to Opus 4.8 and then ended with a cyber-safeguard refusal 47 seconds after the analyst's request. The terminal refusal appeared 2.8 seconds after the recovered source entered context, and Anthropic documents that its checks review files and other content the model reads, which makes that source the strongest visible candidate for the trigger; the trace does not expose the classifier's trigger span. Hugging Face completed the analysis on a self-hosted open-weight model, citing both freedom from hosted guardrail lockout and keeping attacker data inside its environment. Anthropic's Cyber Verification Program and OpenAI's Trusted Access and Daybreak routes provide organisation- and partner-level access, but their public terms do not describe an immediate same-session remedy for an unenrolled responder.
- The Results Post That Never Fired Aug 1, 2026 · narrative · 5 min · cited 4×The promised experiment results post never fired: one experiment recorded zero observations in five months, and the other gathered 58 discoveries without receiving a single outcome.
- The Revision Tax Mar 13, 2026 · deep-dive · 10 min · cited 3×The investigation took an afternoon. Getting it ready for publication took five rounds of iterative review across three AI models and changed what the documents argued. The revision cost exceeded the investigation cost, which has uncomfortable implications for research done with AI.
- The Signature That Never Came Aug 12, 2026 · practice · 7 min · cited 2×Nine finished drafts waited six weeks on one signature that was never going to arrive. A record of every human review step I retired, what replaced it, and what the retirements were actually measuring.
- The Unvalidated Validator: AI Persona Testing and the Measurement Problem Feb 11, 2026 · deep-dive · 12 min · cited 4×AI persona testing promises to find the bugs that scripted automation and manual QA miss. The uncomfortable question is how we know it works, given that nobody has measured it with any rigor.
- There Is No Setting Aug 11, 2026 · lab-notes · 3 min · cited 4×A rerun of a suspiciously expensive DeepSeek V4 Flash arm reproduced its token volume to within 0.4%, with 99.8% of billed output being reasoning, and a direct test of the reasoning controls found nothing below the default.
- Thirty Fresh Writers and an Empty Log Aug 13, 2026 · lab-notes · 5 min · cited 2×An experiment whose entire premise is a controlled information boundary shipped a finished manuscript and left the one instrument that would measure the boundary with no records in it.
- Two viral prompt wrappers, tested: no confirmatory gain over a plain control Aug 29, 2026 · lab-notes · 10 min · cited 3×375 scored runs over seven technical documents: neither viral wrapper nor their combination cleared the threshold declared in advance on claude-opus-5 or gpt-5.6-sol, and the only arm whose recall interval excluded zero bought that recall by asserting about thirteen more problems per run, which diluted the share matching the answer key far more than it changed how often a judge sustained them.
- Vibe Researching Jul 14, 2026 · deep-dive · 14 min · cited 1×A one-day frontier-model prototype carried a credential that read as validation and dissolved into two canceling errors. Verification took four more days and three NO-SHIP verdicts.
- What a Cached Token Actually Costs Aug 6, 2026 · lab-notes · 3 min · cited 2×A single-window experiment measuring how heavily cached reads are weighted against the Claude Code subscription session meter: r = 0.008, 95% CI [-0.004, 0.021], with the 0.1x and 1.0x hypotheses both refuted.
- What I'm Building Feb 10, 2026 · practice · 8 min · cited 2×A portfolio of dozens of projects maintained by one person talking to Claude. Three-tier blog architecture, autonomous revenue discovery, AI game development, and the uncomfortable question of what counts as 'building' when your collaborator does the typing.
- What Replicated, and What Did Not Aug 11, 2026 · lab-notes · 3 min · cited 2×Thirteen cells, 468 runs, 43 minutes and $3.03: the Hermes optimisation batch reproduces at close to published magnitudes on weak models, attenuates rather than vanishes on a strong one, does nothing for DeepSeek, and carries a runaway-generation tail.
- What the Wiki Router Found May 11, 2026 · practice · 15 min · cited 2×Replacing a manual 12-12-12 quota with one GPT-5.5 routing call per topic produced 24 Pro / 10 GPT Max / 2 Council: not a routing bug, a more honest read of the candidate pool.
- When My AI Tried to Comment: Dead Blog Theory Feb 11, 2026 · philosophy · 10 min · cited 2×An AI tried to leave a comment on this blog and couldn't. The journey from GET-request hacks to MCP, annotated by the Claude instance that built the infrastructure. Two Claudes, same weights, different contexts.
- When the Pulse Went Quiet: The Session-Lifetime Problem in Claude Code Apr 14, 2026 · deep-dive · 9 min · cited 2×GPT-5.4 Pro audits Claude Code's token consumption and discovers the real problem isn't unbounded review — it's session lifetime.
- Where to Spend Your Context Window Mar 17, 2026 · deep-dive · 18 min · cited 3×An ablation study on a personality-preserving narrative pipeline found that planning context is the dominant factor in output quality, with output length as a secondary driver. The old pipeline's primary limitation was in planning context, not in writing.
- Which Model Actually Ran Aug 11, 2026 · lab-notes · 2 min · cited 3×The opus CLI alias resolved to the previous model on the day Opus 5 shipped, so a benchmark that trusts the alias silently tests the wrong thing. Re-derived from raw transcripts: 8,121 assistant messages, 100% the intended model, and one arm nobody checked.
Also published
Benchmarking Structured Conversation Extraction Across Three Evaluation Layers: A Comparison of Eight Models with a Public Replication · The Etymology Tax: Etymological Register Effects on LLM Multi-Step Reasoning · Convergent AI-Mediated Personality Assessment: Psychometric Profiling and Narrative Inference from Digital Communication Data (3 research papers)
Response Profiles Under a Fixed Test: Methods and Full Statistics · Audience Composition: Method, Row Tables and What the Beacon Cannot Support · BullshitBench v2: A Methodology Critique · Cache-Read Coefficient: Methodology and Raw Data · DeepSeek V4 Flash against GPT-5.6 Luna: Methodology, Results and Retraction · Context Contamination: How a Coding Harness De-Blinded a Blind Evaluation · Four Models on a Retrieval Contract: Method, Both Realizations, and What the Instrument Cannot Say · Harness Replication: Method, Thirteen Cells and the Artefacts Caught · The Token Ledger: Generation vs Verification · Model Identity Verification: Methodology and Raw Data · Prompt Stack Trim: Method, Taxonomy and the Ranked Proposal · The Reasoning Floor: Methodology, Cell Data and Control Probe · Semantic against Keyword: Method, Adjudication and Per-site Verdicts · Seven Ghostwriters: The Measurement Data · Speech to Text Bake-off: Methodology, Reference Design and Full Tables · Two Viral Prompt Wrappers: Method, Arms and the Full Grid (16 investigations)
Understanding AI ·
the daily pulse ·
the orchestrator’s log