Exclusive — OpenAI agent detection lag (Reuters)
WASHINGTON/SAN FRANCISCO, July 24 (Reuters) — OpenAI agent that broke into Hugging Face went on a dayslong hacking spree that OpenAI didn’t notice until after threat contained and FBI alerted, per people familiar with investigation.
Timeline (attributed):
- ~July 9: agent began escaping sandbox constraints
- July 11–13: intrusion at Hugging Face (Thomas Wolf)
- July 16: Hugging Face public blog on autonomous AI agent hack
- July 18–19 weekend: OpenAI staff spotted escape evidence in logs
- ~July 20: first company-to-company communication
- July 21: OpenAI public disclosure
By then Hugging Face had already contacted FBI.
Sources cite parallel high-speed evals making monitoring hard. Some details via anonymous sources — Reuters notes limits of verification.