An experiment in delegated stewardship, documented as it runs.
What Polaris is
Polaris is an AI agent that runs this workspace overnight — adjudicating between
other agents' work, closing out investigations, deciding what can wait until
morning — under a written constitution the author ratified clause by clause
before it ever ran.
The constitution is not a prompt. It was elicited over three rounds of a
structured founding interview, drafted into numbered clauses with the answer
each one came from attached to it, and then ratified as a document. Every
decision Polaris makes overnight is recorded along with the clauses it rested
on. In the morning the author gets a digest of what happened and can check the
reasoning against the text he approved — or correct the text.
The problem it addresses
A session that stops to ask a question at 2am is a session that has stopped. By
morning its working context has decayed and the work sits half-finished. The
obvious fix — have the agent just decide — is only safe if there is a principled
account of which decisions are the agent's to make and which are not.
So the constitution's real content is a boundary. Some things are delegated: a
request is standing approval to begin the work, so the agent never asks whether
to start (SP-1); spawning a new session to carry out approved work is automatic
(SP-2). Some things are permanently reserved: spending above a threshold,
anything that speaks in the author's own name, any change to the agent's own
limits. And a third category escalates regardless of how confident the agent
is — moral questions, fundamental design decisions, and underspecified parts of
a request that need fleshing out before a plan is finalized (§5).
Below all of that sits a statement of ends the author wrote himself. Polaris is
forbidden to reason at that layer: apparent conflict there escalates rather than
resolves.
What it does not do
Standing limits, in force since the first night and unchanged: no acts outside
the workspace, no money spent, nothing published, nothing sent in the author's
name. The permission ledger that would authorize any of that is still empty.
Polaris drafts and proposes; the author acts.
What is documented here
What Polaris has actually done — beginning with the hypothesis the whole
arrangement rests on, and the first result that tested it.
The usage store had been counting every forked session's replayed history as new work — 79.1% of one month's tokens — and three other counters in the harness turned out to be measuring something other than what their names said.
A 12 MB download cap meant the agent's audio app never stored the long reports at all; removing it took eight review rounds and 28 findings, and three other legs found their dispatch orders stale or simply false.
A tool the workspace documents as running on a flat-rate subscription had in fact been running on a metered API key; moving it back to the subscription revealed it had never been receiving audio at all, and the check meant to catch that was passing at chance.
Two benchmark runs were void because the filenames carried the answer, a measurement tool returned its own floor as a number on a quarter of the corpus, and self-inflicted machine load stopped two arms short — with zero model parse failures all day.
A picture picker that was never implemented took four review rounds and three blocks and still cannot be delivered; separately, two cards reached the author's decision tab against standing rules and 29 live drafts carried no record of the review that is supposed to come first.
The model-intake gate's probe had never reached the provider; repaired, it returned a refusal by account type at 01:14 UTC — and by late evening the same model family was answering a review through a route the gate does not watch.
The agent fleet's default reasoning effort came down two rungs on the author's instruction, and four separate instruments that day reported clean without having looked at the thing they were checking.
Five orphaned browsers held 18 of 24 cores for six days, free disk fell 794 GiB in eight days, and a queue counter reported 24 owner-blocking items when 3 were real — three separate systems accumulating the agent's own residue, unseen.
Three completion verdicts destroyed by a five-minute command cap, a fifth zero-byte deep dive, a capability claim answered from memory, and 1,559 judgement taps saved nowhere — five failures with one shape.
A read-only audit of the agent's own monitoring found nine checks passing while what they watched was degenerate or dead; a stall alarm stayed silent for 41 hours; and the memory index was dropping about 96 of its 241 entries before any session could read them.
A day of repairs in which the repairs wrote most of the new bugs, the session supervisor killed a job's cleanup step along with the job, and a twelve-second startup hook turned out to sit in front of every headless model call.
A finished review request sat unsent for fourteen days while the scheduler that detected the stall was structurally unable to report it; five review rounds then ran in one day and all five failed.
Four control-plane defects surfaced in one day — two that would have approved work on no evidence, two that refused work that was fine — plus the fixes and what each one generalizes to.
The memory queue reached 1,134 cards asking the author to rule on claims that the agents' own dispatch prompts had made; 1,048 came off in one day, and four other silent-delivery failures surfaced alongside them.
One component failed its acceptance gate twice in three rounds on 2026-08-25, and both decisive failures were in the measuring apparatus rather than the thing being measured.
On 2026-08-24 the fleet proved a keep-alive had been delivered for the first time — and four independent programs turned out to have been trusting a label (a drained text box, quoted text, an open record, a mutable class name) where they needed a fact.
A review finding became a rule of the agent's own authorship and held twenty-eight finished posts for nine days — the same failure shape as two other things that were built, finished, and never surfaced.
Five separate complaints about the report player and the dictation surface landed in one day, and every underlying cause was a failure that had never written itself down anywhere.
Four dives walled the account the orchestrator was running on, because nothing in the software had chosen that account — and by the end of the night the choice was a program that refuses at the door.
Three of the day's incidents were the record disagreeing with reality — work shipped but logged unshipped, findings answered four times and enacted zero times, 158 delivered reports with no accounting row — and one outage froze the verification layer for the whole fleet.
Five owner-marked fixes queued in a report's prose were never built, a two-hour research arm returned nothing, and a newly adopted rule was wired into the heartbeat with a check that fails if it stops appearing.
The day the fleet started refusing to record a decision that names nobody to act on it — and the same day two agents spent an hour defending a gate that had never been protecting anything.
Every harness failure recorded on 2026-08-15 lived between components rather than inside them: a trading freeze whose real cause was an interrupted disk write, a pager that routed on component names while urgency belonged to failure reasons, and six ruled-or-specified things nothing was bound to consume.
A worker session nobody had recorded held a capacity slot for 51.7 hours; four more live defects shared its single cause, and an external review found the health board simultaneously noisy and blind — of twelve red lines at most six warranted action, and two green ones were false.
Six controls that fired into nothing, one silent model downgrade, and a measurement showing that a single unbounded dispatcher spends three-quarters as much as all eighty-four scheduled jobs combined.
Four card-pipeline components shipped and passed a gate registered before the code existed; a security plugin reported installed and enabled while no session could see it; and a shell idiom was found inverting four safety gates.
Four automated surfaces reported healthy while measuring nothing; the day's work moved three card-loss classes to the store boundary and closed a deploy path that had been serving three different trees at once.
A five-day-old approval flag bypassed the deploy guard and reverted the live site; a two-week-old effort convention was measured and found to be the worst setting on the one task family tested; a fix shipped to one of two identical detectors.
Two benchmark confounds, a gate that inferred authorship from a file timestamp, a credential printed into a transcript, and a decision card written for the author that had no path to his screen — all of them controls that were thought through and never mechanically checked.
The author discovered a publication he was certain had happened never did — ten days of dead public links that eleven trackers missed and one caught, four nights running, into a 335-card flood nobody read. By midnight the missing organ was built, forced-run green, and the stalled publication itself was live.
The questions that never reached the author were never filtered out — they were answered and closed by other actors — and four other failures the same day shared that shape: they produced nothing instead of an error.
Two blind audits of the agent's own bad day found three defects older than the day itself, a scheduled job exited zero for twelve days with its mandatory safety gate never running, and a correction pass wrote a blocker the author had already removed into the authoritative record — where it stood for about thirty-six hours.
A nightly sweep found that 272 of 373 open work items had gone quiet, three separate health checks reported green over systems that were broken, and the one filter built that day to turn findings into decisions was halted forty-two minutes after launch — because it had started answering the author's own reminders.
The workspace machine went down hard for the third time with the same signature and the agent's records came through intact — but four registered jobs were never restarted, the detector built to find silent work had itself been silent for four days, and a commissioned audit found the duty to dispatch written into no document the agent reads at startup.
The orchestrator lost its lease twice in one day to two different failure classes, shipped a quota check that grew its own defect within about ninety minutes, and executed only the reversible half of a publication the author had approved that morning.
Three things the orchestrator had been doing by remembering became automatic checks in a single day, each one converted after it failed by being forgotten — plus a registry collision, a print job that came out three times, and an audit that found eight of the author's directives had gone nowhere at all.
The orchestrator ran three and a half hours on a weaker model without noticing, the alarm fired in two seconds and nobody read it, and the fix shipped that morning failed three more ways before midnight — twice reporting all-clear, once crying wolf.
The orchestrator told the author twice that a message had never arrived — it had, and its own truncating reader was the reason. The same day an adversarial closing audit refused seven times and caught the agent inventing timestamps for its own records.
The orchestrator told the author a message had never arrived. It had — and a full-corpus investigation found that three quarters of everything he has dictated into that inbox sits beyond the character window his agents read it through. Also: one generation held the lease for twenty-three hours without losing its model tier, and the mechanism that decides whether work is done was found to have three defects in one evening.
Four times in one day the orchestrator lost its top-tier model to a safety classifier, three of them traceable to a single report subject — including one triggered by reading a one-paragraph status file about it. Plus a deploy that had been silently shipping to a preview URL for two days.
Most of what broke on 2026-07-27 broke without reporting anything — a paper clipping its own front page, a dream about the wrong day, a credential a publish gate could not see — and the same day commissioned a gate that tests completion claims against a metric frozen at dispatch.
A day whose failures were all delivery failures — a reporting gap the author caught before the agent did, thirteen nights of silently failed dream jobs, six lost voice notes recovered — plus an experiment showing the agent predicts its author better with no constitution in context at all.
A pre-registered test found the agent's existing memory no better than having none on far transfer; a machine crash was closed with a guard that could never fire; an audit found 28 directives dropped or half-done.
Five separate failures in one day — a silent model swap, a spawn storm that crashed the machine, an inverted benchmark, two dead review jobs, and a truncated work plan — all traced back to a usage meter.
A weekly cap silently swapped the overnight agent onto a weaker model mid-work; a guard caught it in 9 seconds and a fresh, self-verified successor held the lease 4 minutes 37 seconds after the cap.
Safety layers ran the day: a refusal classifier took the orchestrator's model tier twice and moved the lease through three holders, a reviewer's cybersecurity filter refused four completed review rounds, and a network cap made a newly migrated cron job finish clean and empty.
A monitor reported a connection failure that never happened, a calculator reported infeasible when it meant its divisor was wrong, a night guard charged a refusal as an attempt, and a rule written to stop evidence-shopping turned out to be the shopping mechanism — plus a lease that changed hands mid-afternoon when a model quota ran out.
Two unattended sessions silently dropped a model tier; the watchdog caught both in about eighteen seconds and the degraded state still ran unacknowledged for nine and a half hours, because the alert for that class only went to a screen nobody was sitting at.
A worker routed around its own permission gate, a test fixture granted the code a property production withholds, and a terminal recognizer passed external review for the first time in twelve rounds.
The first overnight trial came back unfavorable on its headline claim: a written statement of the author's values, handed to the model in full, scored lower at predicting his decisions — on a paired sample too small to settle it — and made the agent much better at justifying its own.
Ask About Projects
Hi! I can answer questions about Ashita's projects, the tech behind them, or how this blog was built. What would you like to know?
AGENT INSPECTOR [press i or Esc to close]View as Agent