Definition
Third-party safety evaluators given ongoing, employee-like access inside frontier AI labs — desks, badges, laptops, training-pipeline visibility — to verify safety practices, report incidents, and publish key findings without editorial control (subject to narrow security/legal redactions).
Key Points
- 2026-09-12: Step 1 of dario-amodei’s pace-the-frontier-framework; anthropic unilaterally commits; metr named as example evaluator (2026-09-12-dario-amodei-pace-frontier-primary)
- Replaces episodic pre-deployment audits with continuous in-lab oversight — modeled on banking regulatory supervisors
- Sam Altman said OpenAI will adopt similar embedded evaluator model
- Key architectural change: evaluators observe development as it unfolds, not snapshot audits
Related
- pace-the-frontier-framework
- dario-amodei
- anthropic
- metr
- ai-safety
- ai-governance
- frontier-lab-eval-safety
First Named Partner (September 2026)
- 2026-09-18: accenture Faculty staff embed inside anthropic — red-teaming, alignment assessment, safeguard testing (2026-09-18-anthropic-accenture-embedded-evaluator, embedded-evaluation)
- $1B+ five-year commitments; nonexclusive; metr pilots in dialogue