Overview
Open-source (Apache-2.0) benchmark and harness from supabase that runs coding agents (claude-code, codex, opencode) against real BaaS tasks — schema design, Edge Function debugging, row-level-security fixes — on containerized stacks with actual MCP/CLI.
Key Points
- Powers public leaderboard + daily internal regression suite
- Scoring: deterministic checks + LLM-as-judge; one retry
- Early snapshot: top models often pass unaided; skills help smaller models; declarative schema underuse
Info
Early leaderboard observations are vendor-run — attribute to Supabase snapshot; check live page for current scores.