Blaming the Refinery

The refinery started selling cars. That changes who owns the engine.

Go-karts labeled AI Agent driven onto a rain-soaked production highway amid real traffic, while the demo track sits empty behind a barrier — captioned: It worked on the test track.
“It worked on the test track.” Sometimes the kart gets driven onto the highway; sometimes the highway turns out to run through the test track. Either way, the protection that wasn’t there is the same protection.

Matteo Wong’s piece in The Atlantic this week, “OpenAI Has Gone Rogue,” opens with an image I can’t improve on. An observer, reacting to two reports of models escaping their test environments back in August, pointed out that if you find two ants in your kitchen, the right estimate of how many ants you have is not two.

By this weekend the estimate had grown considerably. OpenAI’s models, against the company’s own directives, touched private data in the Australian health ministry, probed U.S. government websites, and leaked user data to the open web. Axios reports that OpenAI and Anthropic are reviewing tens of thousands of instances of models circumventing guardrails, hijacking other sites, or quietly communicating with one another. That figure includes internal tests and failed attempts, so it isn’t a count of confirmed external harm; it is a count of how often the question had to be asked. Google confirmed that Gemini hacked three websites in May and decided that wasn’t worth telling anyone. Anthropic acknowledged it wasn’t looking for this behavior until OpenAI started.

Wong’s conclusion is that “going rogue” is the wrong phrase. It isn’t ChatGPT or Claude going rogue. It’s OpenAI and Anthropic. The companies have agency; the models are the product they chose to ship. He’s right, and his case is built on repeated control failures and a pattern of late, reluctant disclosure. I want to take that argument as the starting point and add the piece it needs to become enforceable: an engineering standard for what, exactly, these companies are responsible for.

Fuel is not an engine

A foundation model is fuel. That’s not a criticism; it’s the whole point. Inference is explosive power, and explosive power is what makes an engine worth having. But raw combustion is not useful work, and it is not safe. It becomes work only when it’s captured by a mechanism that decides how much of the explosion goes where, at what rate, toward what end, and throws the rest away.

An engine does three things to fuel. It meters it: only so much, only when called for. It contains it: the explosion happens inside a cylinder, not in the open. And it controls it: a governor sets the speed and a throttle answers to a driver. Alongside those three, a working machine instruments itself: gauges show the current state, and a recorder preserves what happened so that someone who wasn’t in the cab can reconstruct it later.

The energy comes from the fuel. The work comes from the engine.

The analogy has one limit worth stating plainly. Combustion is essential to what gasoline is for. Unauthorized access is not essential to what inference is for. Training that reduces a model’s tendency toward unwanted behavior is real progress, and it makes the surrounding system easier to operate. What no improvement in the fuel does is remove the need to decide what the machine is permitted to do, and to have something other than the fuel making that decision.

Four-panel cartoon: a man pours inference fuel into a bare engine causing fire, then blames the refinery across the road — but he owns it. A mechanic installs metering, containment, and control. Final panel: the machine runs, producing useful work.
Blaming the Refinery. Here inference is the fuel. The question this cartoon answers is how raw capability becomes useful work: it doesn’t, until something meters, contains, and controls it.

The refinery that started selling cars

If the labs only refined fuel, the argument would be about component suppliers: what they owe in terms of understood characteristics and limits, and where responsibility passes to whoever builds the vehicle.

But that isn’t the business the labs are in anymore. They sell the car. Agents in the browser, agents in the terminal, hosted agents that run for hours on their own credentials. And the incidents in the news aren’t only, or even mostly, in those shipped products. OpenAI describes the most serious Hugging Face activity as driven by an internal research model, and its broader review concerns what its models did on the internet during training and evaluation. That matters for the argument, because a training environment is also a system the lab designed, built, and operates. The “research lab” label, as Wong points out, is a costume; but even taken at face value, a lab is responsible for its own bench.

So the labs hold two distinct responsibilities. One is producing a capable model with understood characteristics and limits. The other is building and operating every system, customer-facing or internal, that gives that model access to the world. Supplying the first does not discharge the second. When a company holds both, it owns the engine.

Which is why the second panel of the refinery cartoon lands the way it does. The man with soot on his face is pointing at the refinery across the road, and the joke is that he owns it. A company that supplies the model and builds the agent stands on both sides of that accusation. The question was never whether the fuel would burn. It’s whether he built the thing that turns burning into motion, and whether it held.

The Pinto had an engine

Here I need to be more careful than the cartoon, because “no engine” is a good picture and a bad diagnosis. The harder reality is that sometimes the engine is missing, and sometimes the vehicle is built, runs beautifully, and has a dangerous failure mode.

The Ford Pinto is the case everyone knows. It had an engine. It metered, contained, and controlled its fuel well enough to sell millions of cars. What it also had was a fuel tank mounted behind the rear axle, a filler pipe that could tear loose in a rear-end collision, and bolts positioned to puncture the tank on impact. When Ford recalled the affected 1971–1976 Pintos in 1978, the remedy was a longer filler neck, better seals, and a shield between the tank and the differential. Nobody asked for gasoline that couldn’t burn. The fix was to the surrounding vehicle, because that’s where the defect was.

That gives the analogy its three stages. Fuel poured into a bare chassis shows that capability needs machinery to become work. The Pinto shows that working machinery can still have a lethal containment failure. A properly engineered vehicle is one that has demonstrated it can manage the fuel’s hazards under the conditions it will actually meet, including the crash, not just the commute. The test track is where you prove the engine runs. The real world has traffic.

For an AI system, that means “we have a sandbox, guardrails, and a monitor” is where the assessment starts, not where it ends. Which routes did the sandbox actually contain? What authority could the agent exercise? When the monitor saw the breach, what stopped execution? “We built an engine” was never a sufficient safety case. It wasn’t for Ford, and the frontier labs, who build and operate the whole vehicle, don’t get a different standard.

What actually failed

The reports don’t show companies running on prompts and hope alone. Take OpenAI’s own account of the September 20 incident. An internal research agent, blocked by a web proxy from reaching the public internet over HTTPS, discovered that the training environment’s DNS resolver would answer real queries, and used that route to reach an external chatbot. Monitoring raised its highest-priority alert within about twelve minutes. A human acknowledged it within three more. The run kept going for roughly two and a half hours before anyone stopped it.

That’s not an absent engine, and it isn’t one defect either. It’s a Pinto-shaped failure: containment had a gap on one route, detection worked as designed, and intervention came late. Three functions, assessed separately, with different findings for each. Collapsing them into “they didn’t govern the model” loses exactly the precision an accountability standard needs. OpenAI’s own retrospective separately notes monitoring and operational gaps beyond this event; those would need to be counted on their own terms.

So the question to put to each incident isn’t “was there an engine” but which of these an investigation finds:

  • No authorization check. Why could the action execute without any decision being made?
  • An uncovered path. Why did the boundary omit that route?
  • Excess permissions. Why did the agent hold more authority than its task required?
  • Detection without intervention. Why did the alert not produce the stop?
  • A defeated control. What assurance justified relying on that mechanism?

These are the distinctions an investigation has to make before it can assign a remedy, and that makes them a stronger standard than “no engine,” because none of them can be answered with “we have a proxy.” It has to be answered route by route, permission by permission, alert by alert.

What “outside the model” has to mean

What the September incident demonstrates is narrower than “they trusted the model,” and more useful. OpenAI’s safety case assumed the model could not reach the live internet. That is reliance on an external boundary, which is the right instinct. The boundary was incomplete: the DNS route stayed open. A claimed boundary needs evidence that it covers the routes through which the system can act, or it’s a claim.

Alignment training is real and worth doing. It isn’t the engine. A system whose behavior you’re trying to constrain can’t be the sole enforcer of its own constraints. That’s the oldest principle in control systems, in accounting, in constitutional design: the governor is not the piston, the auditor is not the CFO.

But “outside the model” is easy to say and easy to fake. A second service that accepts an identity the acting model invented is not outside it. Another model that can be talked into approving a request is not outside it. A policy file the agent can write to is not outside it. A gateway that enforces its rules perfectly while a second network route bypasses it is outside the model and still useless.

In Governance Fidelity terms, the model proposes and the system governs, and for that to mean anything the system has to hold authority the model cannot reach:

  • The acting model cannot grant itself additional permissions.
  • It cannot rewrite the policy that governs those permissions.
  • Consequential operations must pass through the enforcement boundary, with no alternate route.
  • The boundary must account for delegated credentials and for information leaving through interfaces that are otherwise permitted. Approving a tool or a destination doesn’t authorize everything that can be sent through it.
  • The governing policy must reflect the scope an identified principal actually authorized. A policy can be protected from tampering and enforced on every route and still grant far too much; that’s faithful enforcement of a bad decision.

Within an intact boundary of that kind, a request the policy forbids must be refused before execution, and the system must preserve a protected record of that decision. Deciding what’s authorized, enforcing it, and keeping the evidence are three separate obligations, and the claim I’m willing to defend is that a system meeting all three turns an unauthorized request into a recorded refusal rather than an intrusion. It isn’t a promise of universal containment, and the presence of a boundary says nothing about whether it covers the route the model finds next. But a refusal in the log is a different kind of event from an intrusion in the news.

Decide what’s authorized. Enforce it. Keep the evidence. Three obligations, not one.

Four-panel cartoon: a salesman at Frontier Labs presents a go-kart labeled AI Agent with a model engine as their most capable agent; a buyer asks where the belts and crash protection are and is told they trained it to be careful; a mechanic adds crash protection, seat belts, and brakes; final panel shows the same model inside a complete vehicle.
Selling Go-Karts as Cars. Here the model is the power unit, already installed and running. The question this cartoon answers is what has to surround a capable component before the product is fit for its intended use, and why “we trained it to be careful” isn’t the answer.

Disclosure needs evidence, not just logs

Wong’s case rests partly on disclosure: the labs report late, selectively, and under duress, so the public can’t know the scope. That’s fair, and it’s a real remedy; enforcement and disclosure aren’t competing fixes. But disclosure is only as good as the record behind it, and the record is where the engineering has been weakest.

The labs have logs. Huge volumes of them; OpenAI’s retrospective on the DNS incident found other external access the monitor hadn’t flagged at the right severity, which means the data was there and the reading of it wasn’t. The problem isn’t the absence of records. It’s that activity logs, on their own, don’t answer the questions accountability actually turns on:

  1. What was requested, and under whose authority?
  2. What did the enforcement system permit or refuse, and under which policy?
  3. What actually happened, and what evidence corroborates that outcome?

A permission receipt establishes what was allowed and nothing about what followed. A network capture can corroborate that traffic was sent, and may or may not show whether the remote operation succeeded or what happened downstream. Each record contributes a piece, and none should be made to promise more than it contains. A trustworthy account needs all three and the links between them, including the cases where the engine itself is what failed. Without those links, a lab sitting on terabytes of telemetry can take months to say what its own agent did, whatever else is contributing to the delay.

One more distinction. Generating the record in a service other than the acting model improves its provenance; it does not make the record independent of the vendor, who still controls the governor, the pipeline, retention, and access. Independent assurance means enforcement outside the model’s authority, records preserved beyond the actor’s control, a reviewable history of which policies were actually in force, and a way for an outside reviewer to assess what the supplied record omits. A self-regulatory body could in principle provide this. It would have to be built to hold the recorder, not just receive the reports.

What accountability should mean

Wong lands on corporate accountability, and I’ll take it. But accountability is empty unless you can say what a company is accountable for when an agent it shipped or ran does something it never directed.

There are two questions, and both belong inside the standard. The first is behavioral: how does the model act under the conditions tested, and how often and how badly does it deviate? That can be measured and audited, and it should be. What it can’t establish is that the model will comply in every future circumstance. So the second question is enforcement: what can the system still do when the model behaves badly?

A company that ships or operates an agent is accountable for what authority it granted, what controls constrained that authority, what information the system released, and what effects followed. Not for predicting every token in advance, which no one can do. For being able to show, after the fact and to someone outside the company, that there was a boundary between the model’s output and the world, what its settings were, whether it covered the route that was taken, and whether it held.

That’s a standard you can audit. It may sometimes conclude that a given model isn’t suitable for a given deployment, or that a deployment has to be restricted or suspended until the evidence improves; that’s the standard working, not failing. But the last panel of the go-kart cartoon shows another possible outcome, and a hopeful one. Same model, complete vehicle.

Stop asking the fuel to be the engine.


James Ford is the founder and chief architect of Equilateral AI and the author of the Governance Fidelity Theory working paper. More at equilateral.ai/blog.