EvelynFan/FaceFormer
A Transformer-based neural network that synthesizes realistic 3D facial motions from speech audio.

Not currently ranked — collecting fresh signals.
star history
FaceFormer is an end-to-end Transformer architecture that autoregressively generates sequences of 3D facial meshes from audio input. Given a neutral face template and raw audio, it produces accurate lip movements and facial expressions. The implementation is in PyTorch and includes pretrained models for VOCASET and BIWI datasets.
Frequently asked
- What is EvelynFan/FaceFormer?
- A Transformer-based neural network that synthesizes realistic 3D facial motions from speech audio.
- Is FaceFormer open source?
- Yes — EvelynFan/FaceFormer is open source, released under the MIT license.
- What language is FaceFormer written in?
- EvelynFan/FaceFormer is primarily written in Python.
- How popular is FaceFormer?
- EvelynFan/FaceFormer has 915 stars on GitHub.
- Where can I find FaceFormer?
- EvelynFan/FaceFormer is on GitHub at https://github.com/EvelynFan/FaceFormer.