michaelzhang-ai/Text2Video
A deep-learning system that synthesizes talking-head videos from text input using a phoneme-pose dictionary and GAN-based generation.

Not currently ranked — collecting fresh signals.
star history
This repository implements a text-driven video synthesis system for talking-head generation published at ICASSP 2022. The method builds a phoneme-pose dictionary and trains a generative adversarial network (GAN) to produce video from interpolated phoneme poses. It requires only a fraction of the training data needed by audio-driven approaches, offering more flexibility and faster preprocessing, training, and inference.
Frequently asked
- What is michaelzhang-ai/Text2Video?
- A deep-learning system that synthesizes talking-head videos from text input using a phoneme-pose dictionary and GAN-based generation.
- Is Text2Video open source?
- Yes — michaelzhang-ai/Text2Video is an open-source project tracked on heatdrop.
- What language is Text2Video written in?
- michaelzhang-ai/Text2Video is primarily written in Python.
- How popular is Text2Video?
- michaelzhang-ai/Text2Video has 437 stars on GitHub.
- Where can I find Text2Video?
- michaelzhang-ai/Text2Video is on GitHub at https://github.com/michaelzhang-ai/Text2Video.