← all repositories

openai/mle-bench

A benchmark suite from OpenAI for measuring AI agent performance on machine learning engineering challenges.

1.6k stars Python LLMOps · EvalAgents
mle-bench
Not currently ranked — collecting fresh signals.
star history

MLE-bench evaluates how well AI agents perform at machine learning engineering tasks by running them through a set of standardized ML competitions. The repository includes the dataset construction code, evaluation logic, and baseline agent implementations. The benchmark measures agent capabilities across different difficulty levels (Low/Medium/High) and tracks performance metrics like accuracy and running time.

Frequently asked

What is openai/mle-bench?
A benchmark suite from OpenAI for measuring AI agent performance on machine learning engineering challenges.
Is mle-bench open source?
Yes — openai/mle-bench is an open-source project tracked on heatdrop.
What language is mle-bench written in?
openai/mle-bench is primarily written in Python.
How popular is mle-bench?
openai/mle-bench has 1.6k stars on GitHub.
Where can I find mle-bench?
openai/mle-bench is on GitHub at https://github.com/openai/mle-bench.

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.