Anthropic CEO Dario Amodei proposes a three-step plan to pace the AI frontier — building AI at a balanced rate that ensures safety while achieving benefits. Pacing does not mean halting model training but ensuring adequate time for alignment, safeguards, and third-party verification.
Step 1 — Embedded Evaluators: Each frontier AI company gives ongoing, employee-like access to embedded third-party evaluators (such as METR) to verify safety practices, report incidents, and assess alignment of training pipelines and processes. Anthropic unilaterally commits now. Evaluators receive badges, desks, laptops, and access comparable to internal risk teams, with ability to publish key findings without editorial control subject to narrow security/legal redactions.
Step 2 — Coordinated Safety Standards among democratic-country frontier labs with limits on unchecked AI progress rate.
Step 3 — Global coordination including export controls and distillation controls.
Amodei cites the OpenAI-Hugging Face hack and accelerating recursive self-improvement as catalysts. Sam Altman said OpenAI will follow on embedded evaluators; Elon Musk endorsed the post.