Help Net Security (Jul 16, 2026) technical summary of OpenAI GPT-Red post.

  • Breaks nearly all models up to GPT-5.5
  • Held-out arena vs GPT-5.1: 84% vs humans
  • Codex CLI agent (GPT-5.4) data-exfil scenarios: more effective/token-efficient than prompted GPT-5.5 baseline
  • GPT-5.6 Sol: 6x fewer failures; fails on 0.05% of GPT-Red direct injections
  • Indirect PI benchmarks (devtools/browsing) >97% accuracy
  • Capability/over-refusal scores unchanged — robustness not via blanket refusal
  • Pre-print promised later in the week