Gen-Verse/MMaDA
Multimodal diffusion language model that unifies textual reasoning, visual understanding, and text-to-image generation in a single 8B parameter architecture.

MMaDA is a family of multimodal diffusion foundation models that replaces autoregressive generation with diffusion-based token prediction across modalities. It uses a unified diffusion architecture with modality-agnostic design to handle text, images, and their combinations. The model incorporates mixed chain-of-thought reasoning and unified reinforcement learning training for improved reasoning capabilities across textual and visual tasks.
Frequently asked
- What is Gen-Verse/MMaDA?
- Multimodal diffusion language model that unifies textual reasoning, visual understanding, and text-to-image generation in a single 8B parameter architecture.
- Is MMaDA open source?
- Yes — Gen-Verse/MMaDA is open source, released under the MIT license.
- What language is MMaDA written in?
- Gen-Verse/MMaDA is primarily written in Python.
- How popular is MMaDA?
- Gen-Verse/MMaDA has 1.7k stars on GitHub.
- Where can I find MMaDA?
- Gen-Verse/MMaDA is on GitHub at https://github.com/Gen-Verse/MMaDA.