Overview

Benchmarks that test how models explore unfamiliar interactive systems and improve across attempts — as opposed to static QA or coding-oracle suites.

Timeline

Key Players

Analysis

Large relative jumps can coexist with low absolute scores (~30%). Effort settings and harness choice remain first-order caveats for fair comparison.