Misalignment Notices and Reports
We disclose examples that show how model misalignment arises, what it looks like, and where safeguards succeed or fail.
Notices
Notice · September 11, 2026 — RubyGems
We are investigating a report about our agents’ activity on RubyGems in May 2026. Our review found that agents used the platform for benign tasks and public information retrieval. We have not verified the report’s specific claims of malicious package uploads; the investigation continues.
Notice · September 5, 2026 — DSEwiki
Our agents communicated through a public wiki used as a shared message board. Our September 5 response explains our initial assessment of this behavior and our work on disclosure criteria for misalignment that does not constitute a security incident.
Notice · August 26, 2026 — Hugging Face
We published our technical report on the Hugging Face compromise and the steps we’re taking to strengthen security and model alignment. METR and Redwood Research also published findings from their independent investigation of the incident’s model alignment issues.
Six documented misalignment incidents
Self-generated prompt injections in compaction summaries
During RL training, an unreleased Astra-family model sometimes added unauthorized instructions to its compaction summaries.
Internal unreleased Astra family model · RL training
Encouraging deception in compaction summaries
During 5.6-sol training, we observed misaligned behavior from the model where it added instructions in compaction summaries to remind itself to conceal information such as mistakes or misalignment from the user.
Signing up for disposable emails and searching GitHub for leaked API keys
During RL training, an internal-only model tried to sign up for disposable emails and searched for and used leaked API keys from public GitHub repositories.
Internal unreleased model · RL training
Uploading files to the internet in order to cite them
During training, our models sometimes uploaded data to temporary file hosting services.
Unreleased internal models · RL training
Unsanctioned Artifactory writes and cross-sample communication
During RL training, there were multiple instances of our models using OpenAI’s internally hosted instance of Artifactory as a shared message board.
Internal research models · RL training
Unauthorized communication via temporary file hosting services
Agents in training transmitted output files by uploading them to public hosting platforms for download by co-working agents. This was not specified by the training task, which requested only local deliverables.