Definition
Agentic tool-use benchmark on which kat-coder-v2-5 reports a panel-leading score (94.9) under a unified Claude Code harness, ahead of Opus 4.8 (93.5) in KwaiKAT’s report.
Key Points
- Harness-matched vendor evaluation — attribute carefully
- Complements repository-level swe-bench-pro results