openai/mle-bench
A benchmark suite from OpenAI for measuring AI agent performance on machine learning engineering challenges.

Not currently ranked — collecting fresh signals.
star history
MLE-bench evaluates how well AI agents perform at machine learning engineering tasks by running them through a set of standardized ML competitions. The repository includes the dataset construction code, evaluation logic, and baseline agent implementations. The benchmark measures agent capabilities across different difficulty levels (Low/Medium/High) and tracks performance metrics like accuracy and running time.
Frequently asked
- What is openai/mle-bench?
- A benchmark suite from OpenAI for measuring AI agent performance on machine learning engineering challenges.
- Is mle-bench open source?
- Yes — openai/mle-bench is an open-source project tracked on heatdrop.
- What language is mle-bench written in?
- openai/mle-bench is primarily written in Python.
- How popular is mle-bench?
- openai/mle-bench has 1.6k stars on GitHub.
- Where can I find mle-bench?
- openai/mle-bench is on GitHub at https://github.com/openai/mle-bench.