Picsart-AI-Research/StreamingT2V
An autoregressive video generation model that creates long, temporally consistent videos from text or image prompts by extending Stable Video Diffusion.

This repository implements StreamingSVD, an enhanced autoregressive technique for text-to-video and image-to-video generation. It extends Stable Video Diffusion (SVD) into a high-quality long video generator capable of producing videos up to 200 frames (8 seconds) with rich motion dynamics. The method maintains temporal consistency throughout the video while aligning closely to input text or image prompts. Part of the StreamingT2V research family published at CVPR 2025.
Frequently asked
- What is Picsart-AI-Research/StreamingT2V?
- An autoregressive video generation model that creates long, temporally consistent videos from text or image prompts by extending Stable Video Diffusion.
- Is StreamingT2V open source?
- Yes — Picsart-AI-Research/StreamingT2V is an open-source project tracked on heatdrop.
- What language is StreamingT2V written in?
- Picsart-AI-Research/StreamingT2V is primarily written in Python.
- How popular is StreamingT2V?
- Picsart-AI-Research/StreamingT2V has 1.6k stars on GitHub.
- Where can I find StreamingT2V?
- Picsart-AI-Research/StreamingT2V is on GitHub at https://github.com/Picsart-AI-Research/StreamingT2V.