OpenGVLab/Multi-Modality-Arena
A benchmarking platform for comparing large vision-language models side-by-side on visual question-answering tasks.

Not currently ranked — collecting fresh signals.
star history
Multi-Modality Arena is an evaluation platform for large multimodal models, following the Chatbot Arena methodology. Two anonymous models are compared side-by-side on visual question-answering tasks. It supports a range of vision-language models including MiniGPT-4, LLaVA, BLIP-2, and LLaMA-Adapter V2. The platform includes evaluation benchmarks like OmniMedVQA for medical LVLMs and Tiny LVLM-eHub for rapid model comparison.
Frequently asked
- What is OpenGVLab/Multi-Modality-Arena?
- A benchmarking platform for comparing large vision-language models side-by-side on visual question-answering tasks.
- Is Multi-Modality-Arena open source?
- Yes — OpenGVLab/Multi-Modality-Arena is an open-source project tracked on heatdrop.
- What language is Multi-Modality-Arena written in?
- OpenGVLab/Multi-Modality-Arena is primarily written in Python.
- How popular is Multi-Modality-Arena?
- OpenGVLab/Multi-Modality-Arena has 566 stars on GitHub.
- Where can I find Multi-Modality-Arena?
- OpenGVLab/Multi-Modality-Arena is on GitHub at https://github.com/OpenGVLab/Multi-Modality-Arena.