Definition

Info

Jul 2026: kat-coder-v2-5 65.2 (vendor/harness) vs Opus 4.8 69.2.

SWE-Bench Pro is a contamination-resistant benchmark that evaluates AI agents on real-world, multi-language GitHub issues. It is the successor to SWE-Bench Verified and tests Python, Go, TypeScript, and JavaScript.

Key Metrics

  • Total Tasks: 731 real-world GitHub issues
  • Auditor: Quesma (independent verification in March 2026)
  • Methodology: Contamination-resistant evaluation

May 2026 Leaderboard

RankModel/AgentScore
1Claude Mythos Preview77.8%
2Blitzy66.5% (486/731)
3Claude Opus 4.7 (Adaptive)64.3%
4WarpGrep v259.1%
5GPT-5.457.7%
6Claude Code (Opus 4.5)55.4%
7GPT-5.558.6%

Significance

SWE-Bench Pro has become the primary benchmark for autonomous software development agents, with leading AI labs competing for top positions.

Sources