Overview

Vendors and labs open-sourcing evaluation harnesses that run coding agents on real product tasks (not only synthetic benches) — e.g. supabase-evals.

Timeline

Key Players

Analysis

Real-stack evals (MCP/CLI/containers) surface schema/RLS/docs-usage gaps that SWE-bench alone misses. Vendor-run leaderboards need attribution.