← all repositories

DAMO-NLP-SG/VideoLLaMA2

A multi-modal LLM that processes video and audio for spatial-temporal reasoning and understanding.

VideoLLaMA2
Not currently ranked — collecting fresh signals.
star history

VideoLLaMA 2 is a video large language model that advances spatial-temporal modeling and audio understanding. It extends LLM capabilities to multi-modal video comprehension by combining visual, audio, and text inputs. The project provides model checkpoints, demo spaces on HuggingFace, and training/inference code for the video-LLM architecture.

Frequently asked

What is DAMO-NLP-SG/VideoLLaMA2?
A multi-modal LLM that processes video and audio for spatial-temporal reasoning and understanding.
Is VideoLLaMA2 open source?
Yes — DAMO-NLP-SG/VideoLLaMA2 is open source, released under the Apache-2.0 license.
What language is VideoLLaMA2 written in?
DAMO-NLP-SG/VideoLLaMA2 is primarily written in Python.
How popular is VideoLLaMA2?
DAMO-NLP-SG/VideoLLaMA2 has 1.3k stars on GitHub.
Where can I find VideoLLaMA2?
DAMO-NLP-SG/VideoLLaMA2 is on GitHub at https://github.com/DAMO-NLP-SG/VideoLLaMA2.

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.