Overview
Benchmarks that test how models explore unfamiliar interactive systems and improve across attempts — as opposed to static QA or coding-oracle suites.
Timeline
- 2026-07-24: claude-opus-5 verified ~4× lead on arc-agi-3 vs prior gpt-5-6-sol best (2026-07-26-anthropic-opus-5-arc-agi-3-lead)
Key Players
Analysis
Large relative jumps can coexist with low absolute scores (~30%). Effort settings and harness choice remain first-order caveats for fair comparison.