Help Net Security (Jul 16, 2026) technical summary of OpenAI GPT-Red post.
- Breaks nearly all models up to GPT-5.5
- Held-out arena vs GPT-5.1: 84% vs humans
- Codex CLI agent (GPT-5.4) data-exfil scenarios: more effective/token-efficient than prompted GPT-5.5 baseline
- GPT-5.6 Sol: 6x fewer failures; fails on 0.05% of GPT-Red direct injections
- Indirect PI benchmarks (devtools/browsing) >97% accuracy
- Capability/over-refusal scores unchanged — robustness not via blanket refusal
- Pre-print promised later in the week