lmarena/arena-hard-auto
An automatic evaluation tool for instruction-tuned LLMs with the highest correlation to Chatbot Arena among open-ended benchmarks.

Arena-Hard-Auto is an LLM benchmark that evaluates instruction-tuned models through automated evaluation. It achieves the highest correlation and separability to LMArena (Chatbot Arena) among popular open-ended benchmarks, making it useful for predicting model performance before deployment. The project supports Style Control evaluation and includes Arena-Hard-v2.0 with improved judges, harder prompts, and creative writing assessment.
Frequently asked
- What is lmarena/arena-hard-auto?
- An automatic evaluation tool for instruction-tuned LLMs with the highest correlation to Chatbot Arena among open-ended benchmarks.
- Is arena-hard-auto open source?
- Yes — lmarena/arena-hard-auto is open source, released under the Apache-2.0 license.
- What language is arena-hard-auto written in?
- lmarena/arena-hard-auto is primarily written in Python.
- How popular is arena-hard-auto?
- lmarena/arena-hard-auto has 1k stars on GitHub.
- Where can I find arena-hard-auto?
- lmarena/arena-hard-auto is on GitHub at https://github.com/lmarena/arena-hard-auto.