deepglint/unicom
UNICOM is a large-scale vision transformer model designed as a visual backbone for multimodal large language models like LLaVA.

Not currently ranked — collecting fresh signals.
star history
The repository provides foundational visual representation models trained at scale using LAION400M and COYO700M datasets. It implements sample-to-cluster contrastive learning to optimize vision encoders, and these models serve as the vision tower in multimodal LLM pipelines such as LLaVA-NeXT with Qwen2.5-7B. Benchmarks demonstrate strong performance across document understanding, chart analysis, and general VQA tasks.
Frequently asked
- What is deepglint/unicom?
- UNICOM is a large-scale vision transformer model designed as a visual backbone for multimodal large language models like LLaVA.
- Is unicom open source?
- Yes — deepglint/unicom is open source, released under the MIT license.
- What language is unicom written in?
- deepglint/unicom is primarily written in Python.
- How popular is unicom?
- deepglint/unicom has 701 stars on GitHub.
- Where can I find unicom?
- deepglint/unicom is on GitHub at https://github.com/deepglint/unicom.