Definition
Info
Jul 2026: kat-coder-v2-5 65.2 (vendor/harness) vs Opus 4.8 69.2.
SWE-Bench Pro is a contamination-resistant benchmark that evaluates AI agents on real-world, multi-language GitHub issues. It is the successor to SWE-Bench Verified and tests Python, Go, TypeScript, and JavaScript.
Key Metrics
- Total Tasks: 731 real-world GitHub issues
- Auditor: Quesma (independent verification in March 2026)
- Methodology: Contamination-resistant evaluation
May 2026 Leaderboard
| Rank | Model/Agent | Score |
|---|---|---|
| 1 | Claude Mythos Preview | 77.8% |
| 2 | Blitzy | 66.5% (486/731) |
| 3 | Claude Opus 4.7 (Adaptive) | 64.3% |
| 4 | WarpGrep v2 | 59.1% |
| 5 | GPT-5.4 | 57.7% |
| 6 | Claude Code (Opus 4.5) | 55.4% |
| 7 | GPT-5.5 | 58.6% |
Significance
SWE-Bench Pro has become the primary benchmark for autonomous software development agents, with leading AI labs competing for top positions.