Claude Evolution System
Self-improving AI development environment. Discovers new capabilities, evaluates them against a scoring framework, and integrates approved tools on cron.
- Autonomous capability discovery pipeline
- 5-criterion evaluation scoring framework
- Auto-integration of approved capabilities
- Multi-model orchestration (Claude, Codex, Gemini)
Activity Timeline
- Agent census: 35 catalogued, name collision found; 15 rationalization findings staged.
35 agent definitions catalogued; workspace codex-researcher found to shadow global version (critical collision). 15 rationalization findings (2 P1, 6 P2, 7 P3). Weekly-insights ran zero novel findings — duplicate digest from 3 prior runs.
- Model tracker updated with Qwen3.8 and SenseNova U1.5 Lite; self-observation tab shipped.
LFM25 registry reflects latest open-weight entries. Self-observation UI (1,116 lines) completed with post-gate bugfixes. Sol-cutover remediation ongoing. 55 evals pending.
- Daily heartbeat recorded; 55 evaluations pending in review queue.
Evolution pipeline degraded but stable. Backlog at normal steady-state size.
- CLAUDE.md trimmed 28% (12,598→9,038 chars); v2.1.235 verified current.
Burn-guard sessions removed 4 stale facts across 19 skills. No new GA model releases. 55 pending evaluations tracked across the system.
- Model reference table verified; Investigation Runner framework specified.
GPT-5.6 family and Gemini 3.6 Flash GA dates and pricing confirmed against live vendor docs. Discord monitoring and Brave/Exa/Pro escalation routing specified as an Investigation Runner agent framework — not yet deployed.
- Heartbeat job templated; bq-058 security fix verified; agent specs drafted.
Daily capability discovery heartbeat job templated with Phase 1 sources defined. GPT Max synthesis bypass fix confirmed with 22/0 contract tests. Investigation Runner and Herald agent specs drafted for Discord topic processing and backlog surfacing.
- 54 pending evaluations queued; 27 agents active; idle flag raised on backlog.
Capability discovery pipeline has 54 items waiting without processing across 5 active sessions. No commits. Agent roster at 27 active with 71 registered skills.
- Cron health: completion stamps built, 2 stuck reports cleared, 34 false alerts suppressed.
M7 completion-stamp mechanism tracks report status without production ledger writes. OAuth-orphaned reports cleared. Night-shift-sentinel alert classifier updated. Templated renderer copy discovery revises distinct session load estimate to ~130/day.
- Cordis investigation: identity primitive confirmed, five defects traced to one root cause.
61 KB report produced. tmux 3.4 libsystemd scopes give verified 1:1 pane identity (36/36). Root cause: identity derived post-acquisition rather than at mint-time. Zombie session bg-c-bq-041 — 42 hours past TTL, 18 orphaned processes — cited as primary evidence.
- 6 pre-publish security findings surfaced; 4 commits enforcing gate behaviors.
Write-hook fail-closed, preflight refusal honesty, and private-checkout wording tightened. Two P1 findings (write-hook fail-open, owner-interest gate) and four additional P2s under active remediation; 52 evaluations pending.
- Burn-week session running; 51 evaluations pending; grok-build flag targeted.
Autonomous executor run logged. Five ordered burn tasks initiated including grok-build flag removal and commission-orphan requeue. Infrastructure sitting at 51 pending evaluations, 27 agents, 71 skills on Claude Code v2.1.228.
- NE-1 and NE-4 calibration experiments frozen; overthinking research commission open.
30-cell Fable ladder executed and manually audited, converging on 10 confirmed defects. Reasoning-effort research comparing Fable, Opus 5, and Sol 5.6 commissioned on inverse-scaling patterns. Evaluation backlog steady at 51 pending items.
- Sweep-card rewritten to plain-language output; interview-mode UI approved.
v2.1.226 current, 50 evaluations pending. Sweep-card now produces narrative output rather than structured dumps. Two investigations queued: queue-burn drain-rate analysis and row-343 deep-dive with concrete options.
- Executor run log committed; Claude Code version check initiated (8-version delta).
Commit 95edb13 records autonomous orchestrator activity for the day. Version audit covers 2.1.217–2.1.224 range since baseline 2.1.216.
- Semantic classification layer for nightly cron jobs went live at 93.1%/86.2% accuracy.
Hybrid GPT-5.6 Luna approach activated for closure-verify and silence-sweep paths; all misclassifications conservative-direction. False-claim audit corrected post 067. Dispatch bug identified: 18 of 146 closed items remain dispatchable.
- 14 defects found (3 critical); interview-mode approved; v2.1.222 live.
Defect hunt identified liveness alarm stuck 10 days, session reaper with empty nominations list running nightly, and investigation runner burning ~$155/month on silent channel since July 24. Root cause: hand-maintained queue registry. Interview-mode design accepted and BUILD authorized.
- Daily executor run logged; 49 evaluations in pipeline.
Autonomous executor run committed (fd02004). Pipeline holds 49 pending evaluations across 27 agents and 71 skills — expected standing load.
- Gate queue 2/4 failed; repo reseat strategy approved.
Gates A and B completed; C (exit 141) and D (exit 3) need triage. Scout cost-gate reworked to per-task AA intelligence index priority. Repository reseat approved for three repos; 49 evaluations pending.
- Polaris gen-24 succession complete; model-scout lands with 134 tests; ask-polaris escalation prototype live.
Model-scout dry run processed 414 catalog records and ranked 32 candidates; cost gate refactored to per-task comparison with fail-closed enforcement. ask-polaris prototype enables background agents to route blocking decisions to orchestrator without freezing the session.
- Model-scout shipped; live run found 3 defects. Model state audit confirmed current.
New tool: 4 configs, 6 modules, 134 tests, 414-record dry run in ~4s. Live run caught Inkling-Small throughput shortfall, DeepSeek V4-Flash availability gap, and credential scope escalation. CONTEMPORARY-MODELS-CHECK passed with zero table updates needed.
- Harrier embedding benchmark rejected; shadow adjudicator constitution-bias debiasing documented.
Fusion embedding approach degrades retrieval quality — verdict: do not adopt. Shadow adjudicator accuracy drops from 67.3% to 61.2% when the governance constitution is loaded; three-step baseline-first procedure now documented to correct the conservative bias.
- Backlog triage complete: system is decision-starved, not capability-starved.
146 of 235 backlog rows triaged; ~37 gated on owner decisions, not model capability. 71 of 91 cross-session asks were pure repetition from session-boundary context loss. Multi-arm benchmark (Haiku/Fable/Sol-single) run on frozen-corpus snapshots with regression controls.
- 52/52 tests passing after term-filter and stamp-field idempotency fixes; owner-interest gate added.
Double-counting in synonym pairs corrected in term-filtering lens. Stamp-field markdown misparse fixed in backfill reopen logic. Owner-interest routing gate reopened 10 archived rejects. 146+ backlog rows triaged into burn (48) and codex-ultra (12) queues.
- 278-change housekeeping push; public gists 22→15; term-scrub gate verified.
R1–R15 routine items plus N2/N4/S3/S2 committed to integration branch. N1 (4 posts) held for content-court review. Contemporary model check confirmed GPT-5.6 Sol/Terra/Luna and Gemini 3.6 Flash all current.
- Security scrub of public mirror: 21 terms removed, live API key recovered from plaintext.
Identity-linkage terms, staging IPs, and Moltbook registration credentials purged from mirror source via 23 substitutions. Plaintext API key recovery triggered the audit. Executor run completed cleanly.
- Observer validation clean; 278 changes staged; housekeeping R1–R15 approved.
All 5 background tasks exited cleanly. Housekeeping approvals unlocked the integration/reattach-approved-backlog branch. Roster: 27 agents, 69 skills, 50 pending evaluations queued.
- Gemini version gap filed; night-shift r4-r8 remediation complete; polling infra operational.
Pipeline state file lagged global config on Gemini model version — discovery filed. Night-shift TOCTOU race and four additional defects fixed with probe validation and passing test suites. Discord investigation runner architecture specified.
- Gemini 3.6 Flash discovered and filed; 48 evaluations pending.
Model confirmed against Google docs July 21: 1M context, $1.50/$7.50/M. Daily heartbeat capability-discovery job configured. Polaris-C (Fable) launched on Account B with full authority stack inherited.
- Discord investigation runner designed; daily executor armed; 2 commits deployed.
Investigation runner fetches Discord channel messages, filters topics, dispatches research, writes structured markdown reports. Daily integration executor armed on codex-exec transport. Model reference table validated against current GA dates.
- Integration tail reattached; 16 technique notes indexed; Gemini 3.5 Pro added to eval pipeline.
50 backlog items processed, 16 techniques added to library index across four sessions. Eight commits covered 2.1.193/195 registry updates, hook-matcher audit, Capframe, and defense-in-depth notes. Gemini 3.5 Pro detected in freshness sweep and queued for evaluation.
- Registry backfilled through v2.1.214; Sonnet 5 entry added.
15-version gap (v2.1.200–v2.1.214) closed in one session. 9 discovery files generated for pipeline evaluation. 46 evaluations remain in the pending queue.
- Model version check completed; no new GA releases since GPT-5.6 Sol on 2026-07-09.
State file current as of 2026-07-16. 46 pending evaluations, 25 agents, 64 skills active.
- Model reference audit completed (68 days stale); 46 evaluations still queued.
Fable 5 added, Opus 4.8 and GPT-5.6 Sol/Terra/Luna registered. Discovery file created for agent/skill references needing model ID updates. Evaluation backlog and pending discovery sweep keep status degraded.
- Investigation Runner and Herald agent specs drafted; publish inversion Phase 1 scoped.
Design day: 8 sessions, 0 commits. Two Discord-integrated agents specified. Decision-making pattern established (AskUserQuestion for all forks). Evaluation loop stalled at 40 pending items.
- Approval queue GC: 2 redundant items rejected; safety model unified.
Hardened approval gate with unified orchestrator safety model. Added backlog test-and-deploy method with Opus rubric and preregistration tooling. 40 items queued for next evaluation cycle.
- GPT-5.6 series (Sol/Terra/Luna) discovered; agent specs drafted for Discord and workspace orchestration.
Model landscape updated with GPT-5.6 variants announced June 26. Agent specs drafted: Investigation Runner for Discord inbox scanning, workspace orchestrator with 5-tier priority matrix. 43 evaluations pending in queue.
- Publish model inversion specified (blocklist→allowlist, pending approval); portfolio cleanup; model version freshness audit; 43 evaluations in queue.
Phase 1 of publish model inversion fully designed but held pending owner approval. Portfolio cleanup shipped in 3 commits: README, private mirror backup, registry path and IP fixes. Model version state file found 6 weeks stale relative to CLAUDE.md reference.
- RC warming bug root-caused and fixed; publish-model allow-list infrastructure built; GPT-5.6 Sol tracked.
ANTHROPIC_BASE_URL export during RC launch was causing proxy 404s on availability checks — stripped from launch script. Allow-list publishing model ready for Phase 2 review. Multi-account isolation verified via CLAUDE_CONFIG_DIR workaround.
- Version registry to v2.1.199; Remote Control and newspaper pipeline restored; ChatGPT Pro MCP hardening investigation opened.
RC was blocked by ANTHROPIC_BASE_URL proxy in launch script — patched. Newspaper crontab failed due to missing logs directory — fixed, job 4 completed. ChatGPT Pro MCP failure modes characterized (timeouts, CDP drops, 20-char truncation). Context window 1M clamp confirmed intentional behavior.
- GPT-5.6 family captured in registry; Investigation Runner agent spec completed; ChatGPT Pro MCP hardened.
GPT-5.6 Sol/Terra/Luna in limited preview, public rollout late July. Gemini 3.5 Flash gained computer-use June 24. Investigation Runner automates Discord research pipeline with deduplication state tracking. Catalog: 62 agents, 63 skills, 37 evaluations pending.
- Investigation Runner agent specified for Discord-to-pipeline automation.
Filters a designated channel for owner research requests and outputs markdown reports and JSON summaries to the investigations directory. State persistence via last_processed_message_id prevents duplicate processing.
- Discord investigation runner infrastructure specified; model verification routine executed.
Inbox state tracking and execution pipeline designed. 35 evaluations queued, 62 agents and 62 skills active in the system.
- v2.1.195 changelog: hook matcher breaking change + CLAUDE_CODE_DISABLE_MOUSE_CLICKS documented.
Hyphens in hook matchers now require exact match or explicit pattern — breaking for fuzzy-matched hooks. WezTerm crash root cause identified as IME. ChatGPT Pro integration documented and operational. 35 pending evaluations queued.
- Version to 2.1.193; 4 capabilities queued; Gemini 3.5 Flash logged as released.
Four capabilities queued for evaluation: Bash/PowerShell routing hardening, response text capture, file autocomplete, memory-pressure auto-reaping. Gemini 3.5 Pro delayed to July 2026. 35 evaluations pending review in pipeline.
- Version audit 2.1.187–2.1.191: 3 new capabilities, 1 critical silent bug, model registry refreshed.
Discovered sandbox.credentials, MCP idle timeout guard, and /rewind command. Documented silent hook matcher bug fixed in 2.1.187. Model state file updated after 47-day staleness: Fable 5 added, Opus 4.8 corrected, Gemini 3.5 Flash confirmed GA.
- Three new agent specs drafted: Investigation Runner, Discord Inbox Scanner, Workspace Orchestrator.
All three specifications are complete but unimplemented. 32 pending evaluations in the evolution pipeline continue to accumulate. Heavy session load (7 sessions) with no commits.
- Sofya MCP benchmark complete: 80% URL-ranking, 60% snippet-surfacing — registered non-primary behind routing gates.
42-query evaluation (Tracks 1-2) placed Sofya below Exa on snippet quality and agent backend performance (79.2% vs 83.3%). Investigation Runner agent specified for Discord scanning workflow. 31 evaluations still queued.
- v2.1.185 reviewed, Gemini 3.5 Flash added to registry, Sofya.co benchmarked at 100% success.
40 sessions of monitoring. Model audit identified Gemini 3.5 Flash (GA May 19) and flagged 3.5 Pro for June. 42-query search benchmark compared Sofya, Parallel-Search MCP, and Brave. System at 62 agents, 62 skills, 31 pending evaluations.
- Model reference audit confirmed current; Gemini 3.5 Pro flagged as watch item.
GPT-5.5 and Gemini 3.5 Flash verified against 2026-06-10 corpus audit. Gemini 3.5 Pro predicted GA by June 30 — flagged for next verification pass. Standing state unchanged: 30 pending evaluations, 62 agents, 62 skills.
- v2.1.182 and v2.1.183 investigated; two new agent specs drafted.
Changelog and release cross-reference complete, registry updated, Discord summary posted. Investigation Runner and Discord inbox scanner specs drafted to route research requests into the evaluation pipeline.