zai-org/CogVLM2
Open-source multi-modal LLM combining vision and language understanding based on Llama3-8B.

Not currently ranked — collecting fresh signals.
star history
CogVLM2 is a GPT4V-level open-source multi-modal model that integrates visual and language capabilities. The model supports image understanding and extends to video comprehension through keyframe extraction, handling videos up to 1 minute. It offers multiple deployment options including TGI inference and INT4 quantized versions requiring only 16GB VRAM.
Frequently asked
- What is zai-org/CogVLM2?
- Open-source multi-modal LLM combining vision and language understanding based on Llama3-8B.
- Is CogVLM2 open source?
- Yes — zai-org/CogVLM2 is open source, released under the Apache-2.0 license.
- What language is CogVLM2 written in?
- zai-org/CogVLM2 is primarily written in Python.
- How popular is CogVLM2?
- zai-org/CogVLM2 has 2.4k stars on GitHub.
- Where can I find CogVLM2?
- zai-org/CogVLM2 is on GitHub at https://github.com/zai-org/CogVLM2.