SiliconANGLE (Jul 15, 2026): GPT-Red internal attacker for prompt injection discovery.
Verified claims vs OpenAI primary:
- Self-play RL attacker/defender loop
- 84% vs 13% human arena; 6x fewer direct injection failures for GPT-5.6 Sol
- Fake Chain-of-Thought attacks: >95% on GPT-5.1 → <10% on GPT-5.6
- Not a product; kept internal; precursors used since GPT-5.3
- Limits: weak at multi-turn conversational attacks; limited image-based injection coverage — humans still needed
- Comes weeks after GPT-5.6 release positioning vs Anthropic Claude