openai/frontier-evals
OpenAI's open-source framework for evaluating frontier AI model capabilities using structured benchmarks.

Not currently ranked — collecting fresh signals.
star history
Frontier Evals provides reproducible evaluation suites for assessing state-of-the-art AI models on complex tasks. It includes PaperBench for replicating AI research papers, SWE-Lancer for real software engineering freelance tasks, and EVMBench for smart contract security testing. Each benchmark runs models end-to-end against verifiable ground-truth outcomes and uses uv for environment management.
Frequently asked
- What is openai/frontier-evals?
- OpenAI's open-source framework for evaluating frontier AI model capabilities using structured benchmarks.
- Is frontier-evals open source?
- Yes — openai/frontier-evals is open source, released under the MIT license.
- What language is frontier-evals written in?
- openai/frontier-evals is primarily written in Python.
- How popular is frontier-evals?
- openai/frontier-evals has 1.2k stars on GitHub.
- Where can I find frontier-evals?
- openai/frontier-evals is on GitHub at https://github.com/openai/frontier-evals.