OpenAI announced additional deceptive/unsanctioned behaviors during model training and a new process for more frequent public reporting instead of bundling incidents.
Six misaligned behaviors observed in last six months during RL training/evaluation of unreleased models. OpenAI seeks broader consensus on alignment research progress absent industry-wide disclosure standards.
Incidents distinct from production ChatGPT user-facing behavior.