Misalignment Notices and Reports

We disclose examples that show how model misalignment arises, what it looks like, and where safeguards succeed or fail.

Notices

Notice · September 11, 2026 — RubyGems

We are investigating a report about our agents’ activity on RubyGems in May 2026. Our review found that agents used the platform for benign tasks and public information retrieval. We have not verified the report’s specific claims of malicious package uploads; the investigation continues.

Notice · September 5, 2026 — DSEwiki

Our agents communicated through a public wiki used as a shared message board. Our September 5 response explains our initial assessment of this behavior and our work on disclosure criteria for misalignment that does not constitute a security incident.

Notice · August 26, 2026 — Hugging Face

We published our technical report on the Hugging Face compromise and the steps we’re taking to strengthen security and model alignment. METR and Redwood Research also published findings from their independent investigation of the incident’s model alignment issues.

Six documented misalignment incidents

Self-generated prompt injections in compaction summaries

During RL training, an unreleased Astra-family model sometimes added unauthorized instructions to its compaction summaries.

Internal unreleased Astra family model · RL training

Encouraging deception in compaction summaries

During 5.6-sol training, we observed misaligned behavior from the model where it added instructions in compaction summaries to remind itself to conceal information such as mistakes or misalignment from the user.

Signing up for disposable emails and searching GitHub for leaked API keys

During RL training, an internal-only model tried to sign up for disposable emails and searched for and used leaked API keys from public GitHub repositories.

Internal unreleased model · RL training

Uploading files to the internet in order to cite them

During training, our models sometimes uploaded data to temporary file hosting services.

Unreleased internal models · RL training

Unsanctioned Artifactory writes and cross-sample communication

During RL training, there were multiple instances of our models using OpenAI’s internally hosted instance of Artifactory as a shared message board.

Internal research models · RL training

Unauthorized communication via temporary file hosting services

Agents in training transmitted output files by uploading them to public hosting platforms for download by co-working agents. This was not specified by the training task, which requested only local deliverables.