showlab/Show-o
A single transformer architecture that unifies multimodal understanding and generation by combining LLMs with diffusion models.

Show-o is a research repository presenting a unified multimodal model that handles both comprehension and content generation in one transformer. The architecture integrates large language model capabilities with diffusion-based generation, enabling tasks spanning visual understanding (VQA, captioning) and image synthesis. The work represents advances in multimodal AI by eliminating separate encoder-decoder pipelines in favor of a single unified model.
Frequently asked
- What is showlab/Show-o?
- A single transformer architecture that unifies multimodal understanding and generation by combining LLMs with diffusion models.
- Is Show-o open source?
- Yes — showlab/Show-o is open source, released under the Apache-2.0 license.
- What language is Show-o written in?
- showlab/Show-o is primarily written in Python.
- How popular is Show-o?
- showlab/Show-o has 2k stars on GitHub.
- Where can I find Show-o?
- showlab/Show-o is on GitHub at https://github.com/showlab/Show-o.