← all repositories

ming024/FastSpeech2

A PyTorch implementation of Microsoft's FastSpeech 2 neural text-to-speech model for generating speech audio from text.

FastSpeech2
Not currently ranked — collecting fresh signals.
star history

This repository provides a complete implementation of Microsoft’s FastSpeech 2 architecture, a neural network-based text-to-speech system. It supports multi-speaker synthesis across multiple languages (English, Mandarin) and datasets including LibriTTS and AISHELL-3. The implementation includes training pipelines and inference scripts with support for modern neural vocoders like MelGAN and HiFi-GAN to convert mel-spectrograms to waveform audio.

Frequently asked

What is ming024/FastSpeech2?
A PyTorch implementation of Microsoft's FastSpeech 2 neural text-to-speech model for generating speech audio from text.
Is FastSpeech2 open source?
Yes — ming024/FastSpeech2 is open source, released under the MIT license.
What language is FastSpeech2 written in?
ming024/FastSpeech2 is primarily written in Python.
How popular is FastSpeech2?
ming024/FastSpeech2 has 2.2k stars on GitHub.
Where can I find FastSpeech2?
ming024/FastSpeech2 is on GitHub at https://github.com/ming024/FastSpeech2.

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.