The Remote Labor Index (RLI), jointly developed by the Center for AI Safety and Scale Labs, measures how often AI agents complete real freelance projects at professional quality judged by human evaluators.
July 2026 update:
- Fable 5: 16.1% automation rate (highest ever)
- Opus 4.8: 8.3%
- GPT-5.5: 6.3%
Previous leader Opus 4.6 with Claude Cowork scaffold: 4.17%. At RLI launch eight months ago, best rate was 2.5%. Frontier more than quadrupled in under eight months.
Fable 5: 218 of 240 projects evaluated before U.S. government restricted access. Even if Fable 5 failed all 22 missing projects, rate would be 14.6% — still highest.
Benchmark: 240 projects worth $144,000 combined, sourced from 358 verified freelancers across 3D/CAD, architecture, graphic design, video, audio, data analysis, web apps.
CAIS warns AI judges over-score outputs 2.5–3x vs human evaluators. Human evaluation remains essential for hands-on software use tasks.