This page may contain stale information. Last updated: 2026-05-06
This is a stub page. It needs to be expanded with proper content.
Definition
SWE-Bench is a benchmark that evaluates AI systems on real-world software engineering tasks from GitHub repositories. It tests an AI’s ability to resolve issues and implement features in actual codebases.
Variants
- SWE-Bench Verified: Moderated version with human-verified test cases
- SWE-Bench Pro: Contamination-resistant successor testing Python, Go, TypeScript, JavaScript