SHI-Labs/Versatile-Diffusion
A unified multimodal diffusion framework handling text-to-image, image-to-text, and variation tasks in a single model.

Not currently ranked — collecting fresh signals.
star history
Versatile Diffusion implements the first unified multi-flow multimodal diffusion architecture combining VAE, diffuser, and context encoders to handle multiple generation tasks across modalities. It natively supports cross-modal generation and can be extended to semantic-style disentanglement and dual-guided synthesis. The model uses PyTorch and includes a WebUI for convenient inference.
Frequently asked
- What is SHI-Labs/Versatile-Diffusion?
- A unified multimodal diffusion framework handling text-to-image, image-to-text, and variation tasks in a single model.
- Is Versatile-Diffusion open source?
- Yes — SHI-Labs/Versatile-Diffusion is open source, released under the MIT license.
- What language is Versatile-Diffusion written in?
- SHI-Labs/Versatile-Diffusion is primarily written in Python.
- How popular is Versatile-Diffusion?
- SHI-Labs/Versatile-Diffusion has 1.3k stars on GitHub.
- Where can I find Versatile-Diffusion?
- SHI-Labs/Versatile-Diffusion is on GitHub at https://github.com/SHI-Labs/Versatile-Diffusion.