← all repositories

lmarena/arena-hard-auto

An automatic evaluation tool for instruction-tuned LLMs with the highest correlation to Chatbot Arena among open-ended benchmarks.

1k stars Python LLMOps · Eval
arena-hard-auto
Not currently ranked — collecting fresh signals.
star history

Arena-Hard-Auto is an LLM benchmark that evaluates instruction-tuned models through automated evaluation. It achieves the highest correlation and separability to LMArena (Chatbot Arena) among popular open-ended benchmarks, making it useful for predicting model performance before deployment. The project supports Style Control evaluation and includes Arena-Hard-v2.0 with improved judges, harder prompts, and creative writing assessment.

Frequently asked

What is lmarena/arena-hard-auto?
An automatic evaluation tool for instruction-tuned LLMs with the highest correlation to Chatbot Arena among open-ended benchmarks.
Is arena-hard-auto open source?
Yes — lmarena/arena-hard-auto is open source, released under the Apache-2.0 license.
What language is arena-hard-auto written in?
lmarena/arena-hard-auto is primarily written in Python.
How popular is arena-hard-auto?
lmarena/arena-hard-auto has 1k stars on GitHub.
Where can I find arena-hard-auto?
lmarena/arena-hard-auto is on GitHub at https://github.com/lmarena/arena-hard-auto.

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.