Fantasy-AMAP/fantasy-talking
A diffusion-transformer system that generates realistic talking portrait videos from audio input by synthesizing coherent facial motion.

Not currently ranked — collecting fresh signals.
star history
FantasyTalking produces photorealistic talking head videos driven by audio conditions. It leverages a diffusion transformer architecture (Wan2.1) as the base generative model with Wav2Vec for audio encoding. The system synthesizes coherent facial motions including lip movements, expressions, and head poses to create natural talking portraits. Published at ACM MM 2025.
Frequently asked
- What is Fantasy-AMAP/fantasy-talking?
- A diffusion-transformer system that generates realistic talking portrait videos from audio input by synthesizing coherent facial motion.
- Is fantasy-talking open source?
- Yes — Fantasy-AMAP/fantasy-talking is open source, released under the Apache-2.0 license.
- What language is fantasy-talking written in?
- Fantasy-AMAP/fantasy-talking is primarily written in Python.
- How popular is fantasy-talking?
- Fantasy-AMAP/fantasy-talking has 1.6k stars on GitHub.
- Where can I find fantasy-talking?
- Fantasy-AMAP/fantasy-talking is on GitHub at https://github.com/Fantasy-AMAP/fantasy-talking.