NVlabs/VILA
VILA is a family of open vision language models optimized for video and multi-image understanding tasks.

Not currently ranked — collecting fresh signals.
star history
VILA provides a suite of vision language models designed for efficient multimodal AI across edge, data center, and cloud deployments. The project includes models for video understanding, high-resolution image processing, and long-context video analysis. Recent releases cover OmniVinci for visual-audio joint understanding, LongVILA for million-token context windows, and NVILA for full-stack efficiency optimization of multi-modal model design.
Frequently asked
- What is NVlabs/VILA?
- VILA is a family of open vision language models optimized for video and multi-image understanding tasks.
- Is VILA open source?
- Yes — NVlabs/VILA is open source, released under the Apache-2.0 license.
- What language is VILA written in?
- NVlabs/VILA is primarily written in Python.
- How popular is VILA?
- NVlabs/VILA has 3.8k stars on GitHub.
- Where can I find VILA?
- NVlabs/VILA is on GitHub at https://github.com/NVlabs/VILA.