This page may contain stale information. Last updated: 2026-07-05
Overview
Epoch AI is a research organization studying AI capabilities, compute trends, and benchmark methodology. It co-develops evaluation frameworks with partners including metr.
MirrorCode Benchmark (2026)
- Purpose: Measure autonomous reimplementation of entire software programs from black-box execute-only access + checkable specs
- Scale: 25 target programs spanning Unix utilities, bioinformatics, interpreters, cryptography, compression
- Standout result: Claude Opus 4.7 reimplemented gotree (~16K LoC Go, 40+ commands) in 14 hours for $251; 2,000/2,001 tests passed
- Larger target: Pkl (~61K LoC Apple config language) reimplemented at 99%+ test pass rate
- Benchmark score: Opus 4.7 achieved 56% across full suite (June 2026 full release)
Significance
MirrorCode time horizons (2–17 human-week tasks) substantially exceed METR’s ~12-hour bug-fix horizon for Claude Opus 4.6 — signals capability leap in long-horizon autonomous software engineering.