bird-bench/BIRD-CRITIC-1
A benchmark suite evaluating LLMs on real-world SQL user issue resolution tasks.

Not currently ranked — collecting fresh signals.
star history
BIRD-CRITIC-1.0 is a NeurIPS 2025 benchmark for evaluating large language models on software engineering tasks involving SQL. It focuses on realistic database application issues across SQLite and other engines, providing datasets of user SQL problems, ground-truth solutions, and test cases. The repository includes evaluation code, a leaderboard, and integrates with HuggingFace for dataset hosting.
Frequently asked
- What is bird-bench/BIRD-CRITIC-1?
- A benchmark suite evaluating LLMs on real-world SQL user issue resolution tasks.
- Is BIRD-CRITIC-1 open source?
- Yes — bird-bench/BIRD-CRITIC-1 is open source, released under the MIT license.
- What language is BIRD-CRITIC-1 written in?
- bird-bench/BIRD-CRITIC-1 is primarily written in Python.
- How popular is BIRD-CRITIC-1?
- bird-bench/BIRD-CRITIC-1 has 1.1k stars on GitHub.
- Where can I find BIRD-CRITIC-1?
- bird-bench/BIRD-CRITIC-1 is on GitHub at https://github.com/bird-bench/BIRD-CRITIC-1.