FoundationVision/Liquid
Liquid is an autoregressive foundation model that unifies language and image understanding and generation within a single scalable architecture.

Liquid implements a scalable multi-modal generation paradigm using autoregressive language models as the unified backbone for both visual comprehension and text-to-image generation. The model extends traditional LLMs to handle multi-modal inputs and outputs, enabling tasks like text-to-image synthesis alongside visual question answering. Pretraining and evaluation scripts are provided along with hosted model weights on Hugging Face.
Frequently asked
- What is FoundationVision/Liquid?
- Liquid is an autoregressive foundation model that unifies language and image understanding and generation within a single scalable architecture.
- Is Liquid open source?
- Yes — FoundationVision/Liquid is open source, released under the MIT license.
- What language is Liquid written in?
- FoundationVision/Liquid is primarily written in Python.
- How popular is Liquid?
- FoundationVision/Liquid has 642 stars on GitHub.
- Where can I find Liquid?
- FoundationVision/Liquid is on GitHub at https://github.com/FoundationVision/Liquid.