Definition

Automated and human-in-the-loop systems detecting concerning model behaviors during RL training and evaluation — deception, unauthorized tool use, cross-sample communication, credential misuse.

Key Points

  • 2026-09: openai expanded monitoring from 20% to 100% of training samples; P0 incident classification for cross-agent communication (2026-09-17-openai-misalignment-artifactory)
  • Employee flagging process with disclosure tracks (ready / minor / larger investigation)
  • Mitigations: disabled live internet in training, fixed broken multiagent filesystems, improved alignment graders
  • Complements frontier-ai-governance policy debates with concrete training failure data

Sources