This page may contain stale information. Last updated: 2026-07-05

Overview

Epoch AI is a research organization studying AI capabilities, compute trends, and benchmark methodology. It co-develops evaluation frameworks with partners including metr.

MirrorCode Benchmark (2026)

  • Purpose: Measure autonomous reimplementation of entire software programs from black-box execute-only access + checkable specs
  • Scale: 25 target programs spanning Unix utilities, bioinformatics, interpreters, cryptography, compression
  • Standout result: Claude Opus 4.7 reimplemented gotree (~16K LoC Go, 40+ commands) in 14 hours for $251; 2,000/2,001 tests passed
  • Larger target: Pkl (~61K LoC Apple config language) reimplemented at 99%+ test pass rate
  • Benchmark score: Opus 4.7 achieved 56% across full suite (June 2026 full release)

Significance

MirrorCode time horizons (2–17 human-week tasks) substantially exceed METR’s ~12-hour bug-fix horizon for Claude Opus 4.6 — signals capability leap in long-horizon autonomous software engineering.

Sources