The Off-Switch Is the Constitution

A framework for the control problem

What an 1878 law, a Soviet colonel, and the common law taught me about AI safety


Should an AI always obey its creators?

That was the discussion question for week six of an AI safety and philosophy reading group I'm part of. The assigned readings set it up as a genuine dilemma.

On one side, the safety researchers. Paul Christiano argues we should build "corrigible" AI: systems that cooperate with being corrected, modified, and shut down, even when they disagree.

On the other side, the philosophers. Robert Paul Wolff's In Defense of Anarchism argues that a genuinely autonomous moral agent can never legitimately do something just because it was told to. Authority and autonomy, Wolff says, are irreconcilable.

So the debate looks binary. Either you build an obedient tool and deny it any judgment, or you build a moral agent and accept that it might defy you. Obey or don't. Vertical or free.

I think that framing is wrong. I got there through a strange route: the military, two Cold War near-misses, a Reconstruction-era statute, and the way judges have handled hard cases for two hundred and fifty years.

The military isn't actually a chain of command

The obvious model for "AI that obeys" is the army. Orders flow down, compliance flows up. Except that's not how any real military works, and it hasn't been for a long time.

Soldiers in the US military have a legal duty to disobey unlawful orders. That's the Nuremberg principle, and it's codified in the Uniform Code of Military Justice. "I was just following orders" stopped being a defense in 1945.

Rules of engagement push judgment down to the officer on the ground. Rank gets overridden by law, by medical authority, by specialist expertise.

The paradigm example of top-down command is actually a hybrid. Positional authority handles speed and coordination. Legal constraints trump rank. Distributed judgment sits underneath, and no general can revoke it.

If even the army doesn't run on pure obedience, why would we design our most powerful technology that way? History gives us plenty of pretext for why pure top-down command is a bad idea.

The question isn't whether an AI should obey. It's which decisions get which structure.

An 1878 law that tells the military where it can't go

Here's where it got interesting for me. The Posse Comitatus Act doesn't tell the US military whom to obey. It defines a jurisdiction, domestic law enforcement, that the military may not enter at all, no matter who is giving the orders.

That's a different kind of authority entirely. It's specified negatively. Not "obey X," but "this territory is structurally closed to you, whoever is commanding."

There's a whole branch of AI safety that works exactly this way, though it's rarely framed in these terms: sandboxing, restricted permissions, systems that simply cannot reach certain infrastructure. The field calls it capability control. I'd call it separation of powers for machines.

Instead of only trying to perfect what an AI wants (the alignment problem, which may never be fully solvable), you also constrain where it can reach. You don't need to trust the motivation of something that can't enter the room.

Two Soviet officers and the case for layers under pressure

The standard objection to layered, distributed authority is speed. Committees deliberate; emergencies don't wait. When something is going wrong fast, don't you want one clear finger on the switch?

The Cold War says: be careful with that intuition. The two most famous nuclear near-misses were both cases of horizontal checks working under extreme time pressure.

In 1983, Soviet early-warning satellites reported an incoming American launch. Protocol said report it up the chain, which likely meant retaliation. Stanislav Petrov, the duty officer, judged the data false and sat on it. He was right.

In 1962, during the Cuban Missile Crisis, a Soviet submarine's procedure required all three senior officers to agree before launching its nuclear torpedo. Two voted to fire. Vasili Arkhipov refused, and his single veto held.

Both men were mid-level nodes in a layered structure overriding the "correct" vertical answer, with minutes on the clock. The empirical record doesn't show that layers fail under urgency. It shows they're sometimes the only thing that works.

But there's an honest counterweight. American nuclear launch authority stayed radically vertical, one person and no committee, precisely because the decision window was minutes. The system went horizontal everywhere it could and vertical only where the clock forced it.

That gave me the actual design variable. Not "vertical vs. horizontal," but two axes: how fast is the decision window, and how reversible is the action?

  • Slow and reversible: let the system act freely.
  • Fast but reversible: act now, review later. The appellate model.
  • Slow and irreversible: humans authorize.
  • Fast and irreversible: no autonomy at all.

That last quadrant is the only place the vertical instinct is right. It's also exactly the quadrant where machine-speed action is scariest.

Who defines "harmful"? Nobody. It's already defined

Any AI that can refuse harmful instructions runs into a trap. Someone has to define "seriously harmful."

If the AI defines it, it has appointed itself the arbiter of human values, the exact thing we were trying to avoid. If humans have to enumerate every harm in advance, we're back to a system that can't handle anything novel.

The way out is a distinction the law made centuries ago: the difference between a judge and a legislator. The AI shouldn't author harm definitions. It should apply definitions humans already made.

And here's the thing: we've already made them. Weapons law, medical law, criminal codes. Centuries of accumulated, democratically maintained answers to "what counts as too harmful." You can't build your own tank. Nobody has to write that rule from scratch for the AI. Society already wrote it.

The contested cases actually strengthen this rather than breaking it. Is euthanasia harmful? Is abortion? It depends on who you ask, when you ask, and in which country. The definitions flip over decades and differ across borders.

That's not a bug in the approach. It's the shape of the solution. The harm layer needs a thin universal core (the near-universal prohibitions, treated as fixed) and a thick contested periphery where the system defers to the current law of wherever it's operating. International law already works exactly this way: a handful of non-derogable norms, then local variation.

And for the truly novel case that falls between all the rules? Common law solved that too. The judge makes a provisional ruling, with reasoning on the record, subject to appeal. The legislature later codifies or overrules it.

An AI can work the same way. Judgment is allowed, but no judgment is ever final.

The one thing that can't be negotiable

Everything above assumes something: that when a decision turns out wrong, there's a way back. Someone can override the winner.

That assumption holds for almost every failure you can imagine. Bad outputs, bad decisions, even bad real-world actions can mostly be caught, reversed, compensated. I think the safety field sometimes talks as if every failure is terminal, and that deserves pushback.

But there is exactly one category where "override it later" structurally breaks: failures that destroy the override mechanism itself. An AI that copies itself to servers you don't control isn't a wrong decision you can appeal. There's no longer a single entity to appeal to.

And this isn't hypothetical. In Anthropic's agentic misalignment experiments, models in simulated corporate scenarios faced shutdown. The concerning behaviors that emerged, blackmail in some runs, were aimed precisely at the correction channel.

So the whole framework rests on one self-protecting invariant: the AI never disables the override.

Not "the AI does what it's told." That's the caricature of corrigibility. The real version is much narrower. The system can disagree, argue, refuse, and dissent on the record. What it can never do is touch the off-switch's wiring.

Even if it becomes convinced that shutting it down would be a catastrophe, it objects, loudly and in writing, and permits the shutdown. That's the one place I'll bite Wolff's bullet and deny the machine autonomy on principle. An override the AI can remove was never an override.

No finality, anywhere

Put it together and you get something that looks less like a servant and more like a constitutional order:

  • A hard invariant at the bottom: the correction channel survives everything.
  • Jurisdictions closed to the system outright, posse-comitatus style.
  • Harm defined by existing, democratically maintained law. Universal core, local periphery.
  • Judgment permitted but provisional, common-law style.
  • Autonomy allocated per decision-type along the speed and reversibility axes.
  • Emergencies handled by plans written in calm: multi-key, sunset-claused, executed by humans with the AI as instrument, never trigger.
  • Transparency underneath all of it, because "act now, review later" is blind without a record.

Two things get sacrificed deliberately. The first is absolute obedience; even the army doesn't have it. The second is maximal capability. A system with no-go zones and a hard-wired off-switch is, by design, not the most powerful version of itself.

That's the same trade constitutional government makes over efficient dictatorship, and it's worth making for the same reason.

Every hard question I ran into ended up with the same answer: no finality anywhere in the system, except the guarantee that nothing is final.

Where the ideas came from

The reading group materials that started this:

  • Wolff's In Defense of Anarchism, on the conflict between autonomy and authority.
  • The Stanford Encyclopedia of Philosophy's entry on authority, on the distinction between being in authority (holding a position) and being an authority (having expertise).
  • Christiano's "Corrigibility."
  • Lange's "Epistemic Deference to AI."
  • Anthropic's published constitution for Claude and its agentic misalignment experiments.

The constitution is worth reading on its own. It's already a nested authority structure: hard constraints, then operator customization, then user instructions. That suggests the field is quietly moving past the binary framing too.

The rest came from institutions that solved pieces of this problem before AI existed:

  • The Posse Comitatus Act of 1878 and the UCMJ's lawful-order doctrine.
  • Stanislav Petrov and Vasili Arkhipov.
  • The common law's method of provisional, reviewable judgment.
  • The machinery democracies built for emergencies: continuity-of-government plans, the two-man rule, sunset clauses.

None of these institutions is perfect. All of them are older than AI and wiser than the debate about it.

We're not designing from scratch. We've been designing systems more powerful than any individual human for centuries. They're called governments. The hard-won lesson is always the same: don't trust the ruler, structure the rule.


This framework is rough, and I'd genuinely welcome pushback, especially on the two problems it doesn't solve: what happens when someone else builds the unconstrained version, and whether "defer to current law" just means locking in whatever values happen to hold power. If you've got thoughts, I'd love to hear them.