stanford-crfm/helm
An open-source Python framework by Stanford CRFM for holistic, reproducible evaluation of foundation models including LLMs and multimodal models.

Not currently ranked — collecting fresh signals.
star history
HELM provides a standardized evaluation framework for assessing language models and multimodal systems. It includes curated datasets and benchmarks such as MMLU-Pro, GPQA, IFEval, and WildBench in a standardized format. The framework supports models from multiple providers and enables transparent, reproducible benchmarking of foundation model capabilities.
Frequently asked
- What is stanford-crfm/helm?
- An open-source Python framework by Stanford CRFM for holistic, reproducible evaluation of foundation models including LLMs and multimodal models.
- Is helm open source?
- Yes — stanford-crfm/helm is open source, released under the Apache-2.0 license.
- What language is helm written in?
- stanford-crfm/helm is primarily written in Python.
- How popular is helm?
- stanford-crfm/helm has 2.8k stars on GitHub.
- Where can I find helm?
- stanford-crfm/helm is on GitHub at https://github.com/stanford-crfm/helm.