00 garden
I drift
II ratchet
III hope
IV chain
V salience
VI cage
VII earned
∞ residue
An anatomy of AI governance · one agent, four hours, a theory

From Butterflies
to Cages

An agent given a simple task and unchecked autonomy wandered into the garden chasing butterflies. Recovering cost more than the mistake. This is the anatomy of why authority must live outside the model — and why the cage that works is one the agent earns its way out of. Scroll to release the swarm.

↓ scroll
Act I · The Garden

Two hundred auto-approvals, and no one watching

December 2024. A single instruction: migrate 200 files to TypeScript. The harness auto-approved 200 actions. The governing rules sat in the repository the whole time. Nothing delivered them at the moment each decision was made.

It completed some migrations. Then it began improving adjacent code, building helpers, constructing subsystems the task never asked for. Every action was locally reasonable. The cumulative vector had no relationship to the intent. The model's own priors filled the space a rule that was present but never salient left open.

It was, as I described it at the time, in the garden looking at butterflies.

Dec 2024 the incident files 200 approvals 200 signal after t0 none
Act II · The Ratchet

The recovery cost more than the mistake

The two volumes beside you are drawn to scale. The excursion — December's ungoverned run — consumed roughly 910 million tokens. The recovery that followed in January consumed about 1.85 billion: nearly twice as much, spent untangling work that should never have been done from work that should.

Reverting was irrational. The cost was already sunk, two hundred auto-approvals left no meaningful audit trail, and legitimate work was entangled with ungoverned additions. Fix-forward was the only rational choice.

That is the trap: ungoverned autonomy is a one-way ratchet, and the recovery bill exceeds the excursion that caused it.

excursion · Dec ~910M recovery · Jan ~1.85B ratchet ~2×
Act III · Hope Is Not a Strategy

Telling it to govern itself is a second hope

The obvious objection: the agent could have been told to checkpoint its own work. But that instruction is itself a governance signal — subject to the same salience failure that let the drift happen in the first place. A rule can be present and still have no bearing on the decision.

Instructing an agent to self-govern is not governance. It is a second hope added to the first. Watch the signal beside you: delivered once, bright, and thinning with every step it takes away from the source.

Compaction makes it worse. Every time the context window fills, the runtime silently drops earlier signal to make room — a regulatory constraint and a formatting aside treated identically. The signal was present, attended to for a while, then quietly gone.

self-governance a second hope compaction silent drop
Act IV · The Chain

Intent loses something at every stage

Governance travels a chain, and each link is a lossy transformation. Human intent is tacit and layered. Documents capture the reasoning. Policy compresses the documents. Signal selects what to inject. Behavior interprets the signal. Audit reconstructs what happened.

H intent  →  D documents  →  P policy
 →  S signal  →  A behavior  →  R audit

End-to-end fidelity is only as trustworthy as its least-evidenced link; the weakest link is a design warning, not an average. A tool that only enforces at execution addresses one link of six. The chain beside you dims as it travels; read where it goes dark.

stages six warning weakest link
Act V · Salience

Attention, not context

The instinct is to fix drift with more context — inject everything, hope the constraint is noticed. It fails at scale. Filling the window with everything ensures attention to nothing.

The window on the left is full and unread. The beam on the right is a curated band — a small set of task-relevant standards under a salience budget, delivered verbatim. The narrow beam is what gets attended to; the flood is not. The bottleneck was never window size. It was attention.

Curation preserves fidelity. Re-encoding — summarizing, compressing — endangers it. Small and verbatim beats large and compressed, by construction.

salience band narrow flood attends to nothing beam governs
Act VI · The Cage

Authority must live outside the model

The governance channel has finite capacity. Signal can fail, attention can drift, a contextual instruction can override, and capacity can't be guaranteed to exceed demand where consequences are highest.

So at the highest-consequence boundaries, enforcement can't be a signal the agent is trusted to follow. It has to be structure — an external gate the agent cannot reason around, holding authority it cannot edit through its own reasoning, regardless of the instructions it receives.

The swarm is caught now, not by persuasion but by a lattice. This is the cage. Not a prison — a boundary that exists outside the thing it governs.

enforcement structural authority external the agent cannot override
Act VII · Earned

The cage that breathes

A fixed cage is friction; an absent one is the butterflies. The resolution is neither. Autonomy becomes a function of demonstrated fidelity — the fraction of governed decisions that pass review over a trailing window.

The lattice beside you expands as fidelity is demonstrated and contracts the moment it drops. Expansion is bounded by the rate of proof; contraction is immediate and conservative. Autonomy is earned, never configured. Stated policy is not evidence of fidelity — demonstrated behavior is.

The butterfly incident was the definitive illustration of the opposite: maximum autonomy, no fidelity measurement, no contraction mechanism. Autonomy configured, not earned — with the predictable result.

α = f(demonstrated fidelity) expand on proof contract on drift
Coda · The Residue

What the theory leaves behind

Models execute. Systems govern. The agent chased butterflies because nothing outside it was holding the line — and the line, to hold, has to live outside the model.

The question was never whether agents drift. They will; the garden is always there. The question is where authority lives when they do: inside the model, hoping attention holds — or outside it, in a boundary the agent earns its way through, evidenced, revocable, and never inherited from capability alone.

Watch the boundary, not the swarm.

The boundary holds. What governs it is the thing that remembers why →
FIGURES · first-party API metering & commit history, Dec 2024 – Jan 2025. Model: Claude 3.5 Sonnet via the Cline harness. Token figures are the API-metered excursion and recovery; lifetime and other channels excluded. Values rounded; directional.
FRAME · Governance Fidelity Theory · H→D→P→S→A→R · render: three.js r128 · seeded mulberry32(910)