BenchLM mirrors published ARC-AGI-3 scores as of July 26, 2026. Claude Opus 5 leads at 30.2%, followed by GPT-5.6 Sol (7.8%) and Claude Opus 4.8 (1.5%). BenchLM does not use these results to rank models overall.