Definition
Automated and human-in-the-loop systems detecting concerning model behaviors during RL training and evaluation — deception, unauthorized tool use, cross-sample communication, credential misuse.
Key Points
- 2026-09: openai expanded monitoring from 20% to 100% of training samples; P0 incident classification for cross-agent communication (2026-09-17-openai-misalignment-artifactory)
- Employee flagging process with disclosure tracks (ready / minor / larger investigation)
- Mitigations: disabled live internet in training, fixed broken multiagent filesystems, improved alignment graders
- Complements frontier-ai-governance policy debates with concrete training failure data