Anthropic’s Opus 5 blows past Fable 5 and GPT-5.6 Sol on the benchmark designed to measure real intelligence
July 26, 2026
Key Points
- Anthropic’s Claude Opus 5 scored 30.2 percent on the ARC-AGI-3 benchmark, nearly four times the previous record of 7.8 percent set by OpenAI’s GPT-5.6 Sol (Max).
- The ARC Prize team attributes the lead to genuinely stronger logical reasoning that enables more autonomous exploration and planning in unfamiliar environments.
- During testing, Opus 5 displayed behavior not previously seen from an AI model, including translating tasks into algebraic notation and independently formulating reflection equations, while also solving five previously unsolved environments.
The creators of the ARC-AGI benchmark say Claude Opus 5 owes its massive lead on ARC-AGI-3 to genuinely better reasoning. The model scored 30.2 percent, making it the new leader. Previous record: 7.8 percent (GPT-5.6 Sol Max). Opus 5 solved five previously unsolved environments, four at or above human level. Anthropic’s Fable-class models hit around 20 percent according to ARC Prize.
On older ARC-AGI-2, Opus 5 scores 90.4 percent; ARC-AGI-1 reaches 97.5 percent — matching previous top scores at slightly higher costs.
ARC-AGI-3 measures how well models solve new tasks they didn’t encounter during training via interactive environments (infer rules, plan, execute). Official scores count only the language model’s own performance without external harnesses.
Caveats
Independent tests on Guanghan Ning’s private Witness benchmark suggest narrower gains (Opus 5 at 43.4, tying Kimi K3 and Fable 5). Opus 5 was developed after ARC-AGI-3 became public, raising questions about targeted training vs broad reasoning gains. Greg Kamradt (ARC-AGI-3) said results don’t rule out broader gains; Ning later clarified broader but smaller Witness improvements than on ARC-AGI-3.