fishaudio/Bert-VITS2
Bert-VITS2 is a multilingual text-to-speech model combining VITS2 vocoder architecture with multilingual BERT embeddings for improved voice synthesis.

Not currently ranked — collecting fresh signals.
star history
The project implements VITS2, an end-to-end neural vocoder for TTS, enhanced by multilingual BERT to improve prosody and pronunciation accuracy. Users preprocess training data using webui_preprocess.py, and the system supports multiple languages through BERT embeddings. The architecture builds on prior work from MassTTS and jaywalnut310/vits.
Frequently asked
- What is fishaudio/Bert-VITS2?
- Bert-VITS2 is a multilingual text-to-speech model combining VITS2 vocoder architecture with multilingual BERT embeddings for improved voice synthesis.
- Is Bert-VITS2 open source?
- Yes — fishaudio/Bert-VITS2 is open source, released under the AGPL-3.0 license.
- What language is Bert-VITS2 written in?
- fishaudio/Bert-VITS2 is primarily written in Python.
- How popular is Bert-VITS2?
- fishaudio/Bert-VITS2 has 8.8k stars on GitHub.
- Where can I find Bert-VITS2?
- fishaudio/Bert-VITS2 is on GitHub at https://github.com/fishaudio/Bert-VITS2.