Vchitect/Latte
Latte is a latent diffusion transformer for generating videos from text or image conditions, published in TMLR 2025.

Not currently ranked — collecting fresh signals.
star history
Latte implements a latent diffusion transformer architecture for high-quality video synthesis. The repository provides PyTorch model definitions, pre-trained checkpoints on HuggingFace, and complete training and sampling pipelines. It supports text-to-video and image-to-video generation tasks, serving as the official implementation of the corresponding research paper.
Frequently asked
- What is Vchitect/Latte?
- Latte is a latent diffusion transformer for generating videos from text or image conditions, published in TMLR 2025.
- Is Latte open source?
- Yes — Vchitect/Latte is open source, released under the Apache-2.0 license.
- What language is Latte written in?
- Vchitect/Latte is primarily written in Python.
- How popular is Latte?
- Vchitect/Latte has 1.9k stars on GitHub.
- Where can I find Latte?
- Vchitect/Latte is on GitHub at https://github.com/Vchitect/Latte.