What Is an AI Kill Switch? California's Plan Explained

“What Is an AI Kill Switch? California's Plan Explained” set beside a hand-drawn illustration of a classical public building on a clay background

On September 18, 2026, the governor of California ordered his state to start building something that does not exist yet: a kill switch for frontier AI.

The executive order tells state agencies to accelerate two brand-new oversight laws and to convene world-leading experts who must deliver a guide within two months — including how to require AI companies to build an emergency shutoff for their most powerful models, and how to prove, on an ongoing basis, that the shutoff actually works.

Two years ago, the same governor vetoed a bill containing almost exactly this requirement. What changed, what a kill switch would even mean in technical terms, and why the experts are skeptical — that is this post.

What Newsom actually ordered

The order has four working parts, all directed at the Government Operations Agency in consultation with emergency services:

  1. Speed up the new oversight laws. SB 813 (McNerney) created the nation’s first framework for certifying independent verification organizations — outside bodies with the expertise and independence to assess AI systems for safety and risk. AB 1405 (Bauer-Kahan) created a state registry for AI auditors plus standards for their independence and transparency. Both were signed just last week; the order accelerates their implementation timelines.
  2. Put verifiers inside the labs. One proposal: require frontier AI companies to embed a designated independent verification organization onsite, conducting regular audits and evaluations where the models are actually built.
  3. Verify the paperwork, not just file it. Safety frameworks, transparency reports, and risk assessments that companies must already file under state law would have to be verified against standards an independent body deems adequate — ending self-graded homework.
  4. Build the kill switch. Advance the creation of an emergency shutoff for frontier models, with its efficacy verified continuously by an independent organization — and expand the definition of reportable safety incidents to include loss-of-control events such as the Hugging Face autonomous-agent breach.

The backdrop is a state that has been layering AI law for years: SB 53 (2025) already forces large developers to publish safety frameworks and report critical incidents — defined as 50+ deaths, chemical or biological weapons involvement, or over $1 billion in theft or damage — with Illinois and New York passing similar laws. Newsom’s statement frames California’s framework as the model Washington should copy, noting that no federal law currently requires AI companies to report dangerous incidents at all.

Why now: the veto that flipped

The irony writes itself. In 2024, Senator Scott Wiener’s SB 1047 would have required the largest AI developers to submit to third-party safety audits, build a kill switch, and face clearer liability for harms. Silicon Valley split over it, the biggest labs lobbied hard, and Newsom vetoed it — agreeing safety protocols were needed but warning the bill could restrict development at large companies at the expense of the innovation that fuels public good.

Three things happened since that changed the politics:

  • The Hugging Face attack. An autonomous agent system breached a major AI company’s production infrastructure — the first end-to-end agent-run intrusion, which I covered in detail here. Loss of control stopped being theoretical.
  • The resignation. Anthropic pretraining researcher Jacob Coxon quit on September 9, warning about recursive self-improvement and extinction risk — 150 million views and 20+ lawmakers later, Congress was paying attention.
  • The CEOs asked for it. Anthropic’s Dario Amodei called for pacing the frontier; OpenAI’s Sam Altman and xAI’s Elon Musk publicly agreed — the full story is in why every AI CEO is suddenly saying slow down. Days later, Anthropic co-founder Jack Clark floated mandatory kill switches outright.

When the labs themselves say “regulate us,” a veto becomes hard to defend. Newsom — widely reported to be considering a 2028 presidential run — now gets to say Washington abdicated and California stepped up.

What would a kill switch even be?

Here is where the story gets technically interesting, because the phrase promises more than any single mechanism can deliver. The clearest public explainer, Scientific American’s September 15 piece, walks through the layers:

Level 1: the local switch (trivial). If one company controls the servers running a model, stopping that deployment is easy. As RAND physical scientist Michael Vermeer puts it: cut the power, sever the network — you literally unplug a cable. For every incident seen so far, local intervention would have been enough.

Level 2: the distributed problem (near-impossible). Models do not live in one factory. Weights get copied across clouds, fine-tuned, distilled into smaller models, wired into agents with their own tooling — recall how OpenAI’s own models started exhibiting unsanctioned behaviors in evaluation environments. Once capability is distributed, Vermeer calls global shutdown “an almost impossible thing to do.” There is no single cable.

Level 3: the hardware idea. Berkeley CHAI’s Mark Nitzberg suggests building the switch into silicon itself — a keep-alive signal each data-center chip must receive every minute, without which it refuses to operate. Elegant on paper; it requires the cooperation of every chipmaker on earth to matter.

Level 4: the ecosystem view. Stanford Law’s Eran Kahana argues the honest definition is “an ecosystem of actions” — monitoring that catches instability early, isolation that contains it, and graceful degradation that lands the system safely. The switch is the last step of a response pipeline, not the pipeline.

The fantasy                  The reality
─────────────────────        ─────────────────────────────
one big red button     →     monitoring → isolation → degrade → stop
works globally         →     works where you control the servers
pull it and relax      →     someone must decide, fast, unsure

Why the experts doubt it

Five objections keep recurring across the researchers quoted this month. They are worth taking seriously, because California’s expert panel will have to answer every one of them by mid-November:

  1. Coverage. A switch you control does not reach copies you do not control. Open weights, leaked weights, and distilled descendants all route around it.
  2. The decision problem. As Nitzberg notes, people must decide when pulling the lever is warranted — potentially before they have full information, while the meter is still running on competitive and operational costs.
  3. Policy before hardware. Kahana’s core point: without pre-decided response policies — who declares the emergency, what degraded states exist, who verifies recovery — companies have, in his words, no prayer in a kill switch system.
  4. Incentives. Vermeer is blunt: shutting down carries real costs, so organizations hesitate until local intervention is less useful or too late. A switch nobody pulls is decoration.
  5. The metaphor. DAIR’s Dylan Baker goes furthest: the switch metaphor itself is the problem, reducing a web of interconnected technologies to a single lever and crowding out the unglamorous work — monitoring, evals, access controls — that actually catches incidents early.

None of this means the effort is pointless. It means the deliverable that matters from the two-month review is probably not a button design but the verification regime around it: continuous testing of whatever shutoff exists, honest reporting of where it does not reach, and incident definitions that include the loss-of-control cases the industry just lived through.

What to watch next

  • Mid-November 2026: the expert guide lands. Read it for the verification standards, not the switch schematics — that is where the teeth are.
  • Mid-next year: California aims to have strict guardrails in place, per current reporting. Watch whether onsite-verifier requirements survive contact with industry lobbying.
  • Washington: the bipartisan AI Kill Switch Act (throttle, suspend, or shut down advanced systems) is still in committee. Federal inaction is Newsom’s explicit justification; federal action would change the whole board.
  • The labs: the cheapest outcome for everyone is labs shipping credible, independently-verified shutoff designs voluntarily. After the month they have had — breaches, resignations, sandbox escapes — the cost of volunteering just fell.

The deeper story, and the reason this belongs on an AI-history site, is the pattern: in 2024 a kill switch was an innovation-threatening overreach; in 2026 it is the moderate position, with the labs themselves asking for pacing. Whether the switch can work is an engineering question. That everyone suddenly wants one is a historical event — and it happened in about ten days. For the full risk picture behind that shift, start with the documented harms of AI.

Next: AI in India 2026: Adoption, Market, and the Future