jdh-algo/JoyHallo
JoyHallo is a deep learning model that generates talking face videos from audio input, specifically optimized for Mandarin speech.

Not currently ranked — collecting fresh signals.
star history
JoyHallo is an audio-driven video generation model for creating Mandarin talking head videos. It uses a Chinese wav2vec2 model to extract audio features and employs a semi-decoupled structure to capture relationships among lip, expression, and pose features. The model was trained on 29 hours of Mandarin speech video data collected from JD Health employees, including diverse speaking styles and medical topics.
Frequently asked
- What is jdh-algo/JoyHallo?
- JoyHallo is a deep learning model that generates talking face videos from audio input, specifically optimized for Mandarin speech.
- Is JoyHallo open source?
- Yes — jdh-algo/JoyHallo is open source, released under the MIT license.
- What language is JoyHallo written in?
- jdh-algo/JoyHallo is primarily written in Python.
- How popular is JoyHallo?
- jdh-algo/JoyHallo has 520 stars on GitHub.
- Where can I find JoyHallo?
- jdh-algo/JoyHallo is on GitHub at https://github.com/jdh-algo/JoyHallo.