SiliconANGLE (Jul 15, 2026): GPT-Red internal attacker for prompt injection discovery.

Verified claims vs OpenAI primary:

  • Self-play RL attacker/defender loop
  • 84% vs 13% human arena; 6x fewer direct injection failures for GPT-5.6 Sol
  • Fake Chain-of-Thought attacks: >95% on GPT-5.1 → <10% on GPT-5.6
  • Not a product; kept internal; precursors used since GPT-5.3
  • Limits: weak at multi-turn conversational attacks; limited image-based injection coverage — humans still needed
  • Comes weeks after GPT-5.6 release positioning vs Anthropic Claude