JIA-Lab-research/MGM
A multi-modality vision-language model supporting 2B to 34B parameter LLMs with image understanding, reasoning, and generation capabilities.

Mini-Gemini is a vision-language model framework that extends large language models with multi-modal image understanding and generation capabilities. The framework supports a range of dense and Mixture of Experts (MoE) LLMs from 2B to 34B parameters, enabling image comprehension, visual reasoning, and image generation tasks. Built on the LLaVA architecture, it provides model weights, training code, and inference capabilities through HuggingFace integration.
Frequently asked
- What is JIA-Lab-research/MGM?
- A multi-modality vision-language model supporting 2B to 34B parameter LLMs with image understanding, reasoning, and generation capabilities.
- Is MGM open source?
- Yes — JIA-Lab-research/MGM is open source, released under the Apache-2.0 license.
- What language is MGM written in?
- JIA-Lab-research/MGM is primarily written in Python.
- How popular is MGM?
- JIA-Lab-research/MGM has 3.3k stars on GitHub.
- Where can I find MGM?
- JIA-Lab-research/MGM is on GitHub at https://github.com/JIA-Lab-research/MGM.