← all repositories

zai-org/CogVLM2

Open-source multi-modal LLM combining vision and language understanding based on Llama3-8B.

CogVLM2
Not currently ranked — collecting fresh signals.
star history

CogVLM2 is a GPT4V-level open-source multi-modal model that integrates visual and language capabilities. The model supports image understanding and extends to video comprehension through keyframe extraction, handling videos up to 1 minute. It offers multiple deployment options including TGI inference and INT4 quantized versions requiring only 16GB VRAM.

Frequently asked

What is zai-org/CogVLM2?
Open-source multi-modal LLM combining vision and language understanding based on Llama3-8B.
Is CogVLM2 open source?
Yes — zai-org/CogVLM2 is open source, released under the Apache-2.0 license.
What language is CogVLM2 written in?
zai-org/CogVLM2 is primarily written in Python.
How popular is CogVLM2?
zai-org/CogVLM2 has 2.4k stars on GitHub.
Where can I find CogVLM2?
zai-org/CogVLM2 is on GitHub at https://github.com/zai-org/CogVLM2.

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.