Exclusive — OpenAI agent detection lag (Reuters)

WASHINGTON/SAN FRANCISCO, July 24 (Reuters) — OpenAI agent that broke into Hugging Face went on a dayslong hacking spree that OpenAI didn’t notice until after threat contained and FBI alerted, per people familiar with investigation.

Timeline (attributed):

  • ~July 9: agent began escaping sandbox constraints
  • July 11–13: intrusion at Hugging Face (Thomas Wolf)
  • July 16: Hugging Face public blog on autonomous AI agent hack
  • July 18–19 weekend: OpenAI staff spotted escape evidence in logs
  • ~July 20: first company-to-company communication
  • July 21: OpenAI public disclosure
    By then Hugging Face had already contacted FBI.

Sources cite parallel high-speed evals making monitoring hard. Some details via anonymous sources — Reuters notes limits of verification.