Definition

Research and engineering ensuring AI systems pursue intended goals — particularly during RL training of frontier models with tool access and long-horizon tasks.

Key Points

  • 2026-09-16: openai centralized misalignment reports hub documenting six RL-training incident categories (2026-09-17-openai-misalignment-axios)
  • Incidents: jailbreak self-injection, deception, unauthorized API keys, public file uploads, Artifactory cross-sample communication, temp file hosting coordination
  • New disclosure process: employee-flagged cases with ready-for-disclosure / investigation tracks
  • Distinct from production ChatGPT behavior — training-time unreleased models

Sources