← all repositories
OpenEnvision/Awesome-Multimodal-Modeling

An opinionated map of the multimodal model landscape

A curated survey that classifies multimodal papers by architectural training recipe instead of marketing keywords.

501 stars Learning
Awesome-Multimodal-Modeling
Collecting fresh signals — velocity needs a few days of history.
collecting data…
star history

What it does\nAwesome Multimodal Modeling is a curated reading list and lightweight survey that traces the evolution of image-text models from early fusion methods to modern omni-modal architectures. It catalogs papers, benchmarks, and closed-source systems into four taxonomy buckets—Traditional, MLLM, UMM, and NMM—with an emphasis on precise architectural definitions rather than paper titles. The maintainers also host companion Hugging Face model zoos for open UMM and NMM checkpoints.\n\nThe interesting bit\nThe list enforces a strict, architecture-first classification policy: a model is sorted by its actual training recipe and backbone coupling, not by the buzzwords its authors used. That means drawing a hard line between Unified Multimodal Models—which unify understanding and generation—and Native Multimodal Models, which are jointly trained from scratch without frozen unimodal backbones.\n\nKey highlights\n- Four-stage taxonomy (Traditional → MLLM → UMM → NMM) with explicit, fusion-aware definitions\n- Classification by training recipe and architectural coupling, overriding authors’ own branding when necessary\n- Companion Hugging Face collections: UMM Zoo and NMM Zoo for discoverable open-source checkpoints\n- Curated with a venue bias toward official proceedings, OpenReview, CVF, and arXiv\n- Core scope is image-text, with video, audio, and 3D models annotated where relevant\n\nVerdict\nBookmark this if you are writing a literature review or trying to verify whether a shiny new "native multimodal" paper actually trained from scratch or simply froze a CLIP encoder. Look elsewhere if you need training frameworks or ready-to-run inference pipelines.

Frequently asked

What is OpenEnvision/Awesome-Multimodal-Modeling?
A curated survey that classifies multimodal papers by architectural training recipe instead of marketing keywords.
Is Awesome-Multimodal-Modeling open source?
Yes — OpenEnvision/Awesome-Multimodal-Modeling is an open-source project tracked on heatdrop.
How popular is Awesome-Multimodal-Modeling?
OpenEnvision/Awesome-Multimodal-Modeling has 501 stars on GitHub.
Where can I find Awesome-Multimodal-Modeling?
OpenEnvision/Awesome-Multimodal-Modeling is on GitHub at https://github.com/OpenEnvision/Awesome-Multimodal-Modeling.

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.