VITA-MLLM/VITA
A multimodal LLM enabling real-time vision and speech interaction at GPT-4o-level performance.

Not currently ranked — collecting fresh signals.
star history
VITA-1.5 is an open-source omni-modal large language model designed for real-time vision and speech interaction. It supports bidirectional understanding of video, audio, and text modalities in both English and Chinese. The project provides model weights, inference code, and a technical report, targeting GPT-4o-level conversational and perceptual capabilities.
Frequently asked
- What is VITA-MLLM/VITA?
- A multimodal LLM enabling real-time vision and speech interaction at GPT-4o-level performance.
- Is VITA open source?
- Yes — VITA-MLLM/VITA is an open-source project tracked on heatdrop.
- What language is VITA written in?
- VITA-MLLM/VITA is primarily written in Python.
- How popular is VITA?
- VITA-MLLM/VITA has 2.5k stars on GitHub.
- Where can I find VITA?
- VITA-MLLM/VITA is on GitHub at https://github.com/VITA-MLLM/VITA.