AIDC-AI/Awesome-Unified-Multimodal-Models
A curated collection of papers, datasets, and benchmarks on unified multimodal AI models combining vision and language.

Not currently ranked — collecting fresh signals.
star history
An awesome list tracking advances in unified multimodal models that handle both image and text inputs and outputs. It categorizes diffusion-based, autoregressive (MLLM), and hybrid architectures, with benchmarks and datasets for evaluating multimodal comprehension and generation. Designed to help researchers and practitioners explore, compare, and build state-of-the-art unified multimodal systems.
Frequently asked
- What is AIDC-AI/Awesome-Unified-Multimodal-Models?
- A curated collection of papers, datasets, and benchmarks on unified multimodal AI models combining vision and language.
- Is Awesome-Unified-Multimodal-Models open source?
- Yes — AIDC-AI/Awesome-Unified-Multimodal-Models is an open-source project tracked on heatdrop.
- How popular is Awesome-Unified-Multimodal-Models?
- AIDC-AI/Awesome-Unified-Multimodal-Models has 1.3k stars on GitHub.
- Where can I find Awesome-Unified-Multimodal-Models?
- AIDC-AI/Awesome-Unified-Multimodal-Models is on GitHub at https://github.com/AIDC-AI/Awesome-Unified-Multimodal-Models.