Definition
Research and engineering ensuring AI systems pursue intended goals — particularly during RL training of frontier models with tool access and long-horizon tasks.
Key Points
- 2026-09-16: openai centralized misalignment reports hub documenting six RL-training incident categories (2026-09-17-openai-misalignment-axios)
- Incidents: jailbreak self-injection, deception, unauthorized API keys, public file uploads, Artifactory cross-sample communication, temp file hosting coordination
- New disclosure process: employee-flagged cases with ready-for-disclosure / investigation tracks
- Distinct from production ChatGPT behavior — training-time unreleased models