Supabase Releases Evals: an Open Source Benchmark That Scores Claude Code, Codex and OpenCode on Real Supabase Tasks
- Supabase just dropped Evals, an open-source benchmark that finally forces AI agents like Claude Code, Codex, and OpenCode to stop hallucinating and start building real schemas. No more mock data; these models are debugging Edge Functions and fixing RLS policies in actual containerized stacks. It’s a deterministic reality check for the "AI will replace devs" crowd, scoring them on whether they actually break your database or fix it. While moonboys chase rugs, Supabase is ensuring the infrastructure doesn’t collapse under the weight of lazy code. It’s not a rugpull, but it might be a wake-up call for anyone thinking LLMs can write secure SQL without supervision. Check the leaderboard if you enjoy watching machines fail publicly.