CyberAgentAILab/TANGO
A diffusion model that synthesizes realistic gesture videos from speech audio through hierarchical audio-motion embedding.

Not currently ranked — collecting fresh signals.
star history
TANGO generates co-speech gesture videos by mapping audio features to body motion using hierarchical audio-motion embedding and diffusion interpolation. The model takes speech input and produces corresponding gesture animations, enabling video reenactment with realistic body language synchronized to audio.
Frequently asked
- What is CyberAgentAILab/TANGO?
- A diffusion model that synthesizes realistic gesture videos from speech audio through hierarchical audio-motion embedding.
- Is TANGO open source?
- Yes — CyberAgentAILab/TANGO is an open-source project tracked on heatdrop.
- What language is TANGO written in?
- CyberAgentAILab/TANGO is primarily written in Python.
- How popular is TANGO?
- CyberAgentAILab/TANGO has 1.2k stars on GitHub.
- Where can I find TANGO?
- CyberAgentAILab/TANGO is on GitHub at https://github.com/CyberAgentAILab/TANGO.