jdh-algo/JoyVASA
A diffusion-based method for generating talking portrait and animal videos from audio, producing facial dynamics and head motion.

Not currently ranked — collecting fresh signals.
star history
JoyVASA is a diffusion-based approach for audio-driven facial animation that generates realistic talking heads from audio input. It employs a decoupled facial representation framework with a two-stage pipeline: first extracting disentangled facial representations, then generating facial dynamics and head motion from audio. The method supports both human portraits and animal images, producing natural lip-sync and head movements.
Frequently asked
- What is jdh-algo/JoyVASA?
- A diffusion-based method for generating talking portrait and animal videos from audio, producing facial dynamics and head motion.
- Is JoyVASA open source?
- Yes — jdh-algo/JoyVASA is open source, released under the MIT license.
- What language is JoyVASA written in?
- jdh-algo/JoyVASA is primarily written in Python.
- How popular is JoyVASA?
- jdh-algo/JoyVASA has 874 stars on GitHub.
- Where can I find JoyVASA?
- jdh-algo/JoyVASA is on GitHub at https://github.com/jdh-algo/JoyVASA.