Francis-Rings/StableAvatar
A video diffusion model that generates infinite-length avatar videos from a reference image and audio input.

Not currently ranked — collecting fresh signals.
star history
StableAvatar is an end-to-end video diffusion transformer for synthesizing high-quality, infinite-length avatar videos driven by audio. It takes a reference image and audio as conditioning inputs to generate synchronized talking head videos without post-processing. The model uses a diffusion transformer architecture to handle both visual generation and temporal consistency across long video sequences.
Frequently asked
- What is Francis-Rings/StableAvatar?
- A video diffusion model that generates infinite-length avatar videos from a reference image and audio input.
- Is StableAvatar open source?
- Yes — Francis-Rings/StableAvatar is open source, released under the MIT license.
- What language is StableAvatar written in?
- Francis-Rings/StableAvatar is primarily written in Python.
- How popular is StableAvatar?
- Francis-Rings/StableAvatar has 1.2k stars on GitHub.
- Where can I find StableAvatar?
- Francis-Rings/StableAvatar is on GitHub at https://github.com/Francis-Rings/StableAvatar.