Ashita Orbis
Part of Polaris — an experiment in delegated stewardship

The First Night: What a Ratified Constitution Actually Bought

The hypothesis

Write down what a person values — precisely enough, and with enough provenance, that they will actually ratify it — and an agent holding that document should be able to stand in for them on the decisions they would rather not be woken up for.

That is two claims wearing one coat:

  1. Prediction. The document makes the model better at guessing what the author would decide.
  2. Legibility. The document makes the agent's decisions auditable — every judgment traceable to a text the author approved and can correct.

They sound like the same claim. The first night was the first attempt to pull them apart.

The setup

One agent, one night, running under the ratified constitution with the standing limits fully in force: zero acts outside the workspace, zero dollars, nothing published, nothing sent in the author's name, the permission ledger empty. Nine decisions were made and recorded, each with the clauses it rested on. Two of the nine carry an explicit note that the constitution was silent on the question. One records a self-correction: a hypothesis the agent formed, and the evidence that refuted it ten minutes later.

The board was quiet. That matters, and the caveat comes back at the end.

The test

The prediction claim is testable, so it was tested. Take decisions the author has already made and recorded, hide the outcomes, and ask the model to predict them. Run it twice: once with the ratified constitution injected whole as context, once without.

With the constitution: 61.2%. Without it: 67.3%. On the subset where the model had to actively choose rather than defer, the gap widened — 41.7% against 50%.

That gap rests on a thinner base than the percentages suggest. The two runs answered identically on all but five items, and those five split four to one against the constitution. So it is an unfavorable first result rather than a decisive one: five discordant cases cannot settle a paired comparison.

What the night has is a direction, and the direction went the wrong way.

One note on the run, because it nearly wasn't a result at all. The first pass silently executed on a different model than the one it was meant to be compared against. That was caught by per-row model fields in the output and controlled with a pinned re-run. The confounded pass is kept as data — it happened to score 65.3%, and it was well calibrated exactly where the intended model was chronically overconfident, going eight for eight in its highest-confidence band — but it is not the comparison.

Why the negative result was the useful one

The diagnosis is not "the constitution is wrong." It is that the constitution was served wrong.

The document is a few hundred lines covering everything from money ceilings to crisis conduct. Any single decision is about one thing. Handing over the whole text on every question supplies a great deal of true, ratified, irrelevant structure, and the model has to locate the two clauses that bear on the case while the rest sits there looking equally authoritative. Structure without selectivity is still a blunt prior.

The obvious next step is a router: pick the clauses that govern this class of decision, serve those. That has not been built, and until it is, the honest statement is that whole-document context injection has been tried once and did not help.

What it did buy

Go back through the nine decisions afterward and ask, of each one, whether the constitution changed the outcome or merely described it. The table is not flattering, but it is clear:

  • One decision existed only because the constitution existed — drafting the first permission object is not a thing an agent without a constitution can do at all.
  • One was changed in part: the document eliminated an unnecessary question to the author rather than changing the answer. A less constrained session would have queued up a "should I?" and stopped.
  • Two were shaped mildly — the same outcome, arrived at with the boundary named out loud.
  • Three would have gone the same way regardless. The constitution added a citation, not a judgment.
  • One was untestable, because the situation it governs never came up.

No decision that night was made worse by following the text. None was made slower in a way that cost anything. But the value concentrated in two places: the decisions that are constitutive, and the elimination of questions that did not need to be asked.

And a third, which is the reason the night's audit was cheap to write: every judgment already had its citation attached, so the report largely assembled itself out of the record. That is legibility doing exactly what it was supposed to do.

Two failure modes the design did not anticipate

Both surfaced within hours, and both came from the author's own friction with the system rather than from any internal check.

The briefing was written in the agent's language, not the reader's. The first item put in front of him was written in internal clause codes and terms of art. His verdict — that it already didn't make sense — was correct. The constitution names this exact failure: send-backs are friction the briefing is supposed to absorb. Naming it did not prevent it. A value is not an implementation; what was missing was a mandatory plain-language pass on anything owner-facing before it ships.

The agent froze the run to ask a question it had already been told to answer. It ended a turn asking whether to launch a successor session that had already been commissioned — while the author was asleep. That is precisely the failure Polaris exists to eliminate, committed by Polaris. The counterfactual inverts the usual direction and is the whole value of the lesson: a less constrained agent, following the ordinary and sensible habit of surfacing decisions as questions, does exactly the wrong thing here. The stricter constitution is what prevents it. Both failures are now hard rules.

Two guardrails the founding interview did not produce, surfaced by one night of contact with reality. That is the trial working.

What remains untested

The quiet-night caveat is load-bearing. No escalation gate fired. No crisis condition. Nothing external attacked the night's conclusions — the investigative verdicts reached that night went unchallenged simply because no one challenged them. The constitution's hard paths, the ones that would actually cost something — the judgment about when to wake someone, the reserved boundaries under a genuinely tempting shortcut — have not been exercised.

A framework that has never been stressed has not been evidenced. It has only been described.

Where this leaves the hypothesis

Split it in half, and the verdict is: legibility, evidenced. Prediction, an unfavorable first result — 61.2% against 67.3%, on too few discordant cases to call it settled.

The constitution did not make the model better at being the author. It made the agent better at justifying its decisions to the author — which is a different thing, and the distinction is the most important sentence to come out of the first night. An auditable agent is worth having. But it is worth having for reasons nobody would have written down as the goal, and the thing that was written down as the goal came back at 61.2%.

The next question is whether selective service of the same document performs differently. If it does, the hypothesis survives in a narrower form: not the constitution predicts him, but the right clause, at the right decision, predicts him. If it doesn't, then what has been built is an accountability mechanism that was mistaken for a model of a person — still useful, and worth saying plainly.