NVlabs/OmniVinci
OmniVinci is an NVIDIA research multimodal LLM that jointly processes vision, audio, and language inputs.

Not currently ranked — collecting fresh signals.
star history
OmniVinci is an omni-modal large language model designed to jointly understand vision, audio, and language inputs. It is published at ICLR 2026 and available as a model on HuggingFace. The project includes code, pretrained weights, and training pipelines for this multimodal foundation model.
Frequently asked
- What is NVlabs/OmniVinci?
- OmniVinci is an NVIDIA research multimodal LLM that jointly processes vision, audio, and language inputs.
- Is OmniVinci open source?
- Yes — NVlabs/OmniVinci is open source, released under the Apache-2.0 license.
- What language is OmniVinci written in?
- NVlabs/OmniVinci is primarily written in Python.
- How popular is OmniVinci?
- NVlabs/OmniVinci has 674 stars on GitHub.
- Where can I find OmniVinci?
- NVlabs/OmniVinci is on GitHub at https://github.com/NVlabs/OmniVinci.