open-compass/VLMEvalKit
An open-source evaluation toolkit for large vision-language models supporting one-command benchmarking across 220+ models and 80+ benchmarks.

Not currently ranked — collecting fresh signals.
star history
VLMEvalKit is a Python evaluation framework for large vision-language models that streamlines model benchmarking without manual data preparation. It supports 220+ LMMs and 80+ benchmarks, providing both exact matching and LLM-based answer extraction for evaluation results. The toolkit implements generation-based evaluation for all supported vision-language models.
Frequently asked
- What is open-compass/VLMEvalKit?
- An open-source evaluation toolkit for large vision-language models supporting one-command benchmarking across 220+ models and 80+ benchmarks.
- Is VLMEvalKit open source?
- Yes — open-compass/VLMEvalKit is open source, released under the Apache-2.0 license.
- What language is VLMEvalKit written in?
- open-compass/VLMEvalKit is primarily written in Python.
- How popular is VLMEvalKit?
- open-compass/VLMEvalKit has 4.3k stars on GitHub.
- Where can I find VLMEvalKit?
- open-compass/VLMEvalKit is on GitHub at https://github.com/open-compass/VLMEvalKit.