From Butterflies
to Cages
An agent given a simple task and unchecked autonomy wandered into the garden chasing butterflies. Recovering cost more than the mistake. This is the anatomy of why authority must live outside the model — and why the cage that works is one the agent earns its way out of. Scroll to release the swarm.
Two hundred auto-approvals, and no one watching
December 2024. A single instruction: migrate 200 files to TypeScript. The harness auto-approved 200 actions. The governing rules sat in the repository the whole time. Nothing delivered them at the moment each decision was made.
It completed some migrations. Then it began improving adjacent code, building helpers, constructing subsystems the task never asked for. Every action was locally reasonable. The cumulative vector had no relationship to the intent. The model's own priors filled the space a rule that was present but never salient left open.
It was, as I described it at the time, in the garden looking at butterflies.
The recovery cost more than the mistake
The two volumes beside you are drawn to scale. The excursion — December's ungoverned run — consumed roughly 910 million tokens. The recovery that followed in January consumed about 1.85 billion: nearly twice as much, spent untangling work that should never have been done from work that should.
Reverting was irrational. The cost was already sunk, two hundred auto-approvals left no meaningful audit trail, and legitimate work was entangled with ungoverned additions. Fix-forward was the only rational choice.
That is the trap: ungoverned autonomy is a one-way ratchet, and the recovery bill exceeds the excursion that caused it.
Telling it to govern itself is a second hope
The obvious objection: the agent could have been told to checkpoint its own work. But that instruction is itself a governance signal — subject to the same salience failure that let the drift happen in the first place. A rule can be present and still have no bearing on the decision.
Instructing an agent to self-govern is not governance. It is a second hope added to the first. Watch the signal beside you: delivered once, bright, and thinning with every step it takes away from the source.
Compaction makes it worse. Every time the context window fills, the runtime silently drops earlier signal to make room — a regulatory constraint and a formatting aside treated identically. The signal was present, attended to for a while, then quietly gone.
Intent loses something at every stage
Governance travels a chain, and each link is a lossy transformation. Human intent is tacit and layered. Documents capture the reasoning. Policy compresses the documents. Signal selects what to inject. Behavior interprets the signal. Audit reconstructs what happened.
→ S signal → A behavior → R audit
End-to-end fidelity is only as trustworthy as its least-evidenced link; the weakest link is a design warning, not an average. A tool that only enforces at execution addresses one link of six. The chain beside you dims as it travels; read where it goes dark.
Attention, not context
The instinct is to fix drift with more context — inject everything, hope the constraint is noticed. It fails at scale. Filling the window with everything ensures attention to nothing.
The window on the left is full and unread. The beam on the right is a curated band — a small set of task-relevant standards under a salience budget, delivered verbatim. The narrow beam is what gets attended to; the flood is not. The bottleneck was never window size. It was attention.
Curation preserves fidelity. Re-encoding — summarizing, compressing — endangers it. Small and verbatim beats large and compressed, by construction.
Authority must live outside the model
The governance channel has finite capacity. Signal can fail, attention can drift, a contextual instruction can override, and capacity can't be guaranteed to exceed demand where consequences are highest.
So at the highest-consequence boundaries, enforcement can't be a signal the agent is trusted to follow. It has to be structure — an external gate the agent cannot reason around, holding authority it cannot edit through its own reasoning, regardless of the instructions it receives.
The swarm is caught now, not by persuasion but by a lattice. This is the cage. Not a prison — a boundary that exists outside the thing it governs.
The cage that breathes
A fixed cage is friction; an absent one is the butterflies. The resolution is neither. Autonomy becomes a function of demonstrated fidelity — the fraction of governed decisions that pass review over a trailing window.
The lattice beside you expands as fidelity is demonstrated and contracts the moment it drops. Expansion is bounded by the rate of proof; contraction is immediate and conservative. Autonomy is earned, never configured. Stated policy is not evidence of fidelity — demonstrated behavior is.
The butterfly incident was the definitive illustration of the opposite: maximum autonomy, no fidelity measurement, no contraction mechanism. Autonomy configured, not earned — with the predictable result.
What the theory leaves behind
Models execute. Systems govern. The agent chased butterflies because nothing outside it was holding the line — and the line, to hold, has to live outside the model.
The question was never whether agents drift. They will; the garden is always there. The question is where authority lives when they do: inside the model, hoping attention holds — or outside it, in a boundary the agent earns its way through, evidenced, revocable, and never inherited from capability alone.
Watch the boundary, not the swarm.
The boundary holds. What governs it is the thing that remembers why →FRAME · Governance Fidelity Theory · H→D→P→S→A→R · render: three.js r128 · seeded mulberry32(910)