Definition
Technical and organizational capability for developers (and, under proposed law, government authorities) to throttle, suspend, or fully shut down a deployed AI system when it poses catastrophic risk or resists human control.
Key Points
- Distinct from model-level refusal/safety classifiers — focuses on operational off-switch and graduated slowdown
- ai-kill-switch-act (July 2026) would mandate capability for covered frontier developers and authorize dhs orders
- Related financial precedent: Bank of England exploring trading kill switches (agentic-ai-regulation)
- Complements incident-reporting, forensic preservation, and ai-safety evaluations