NVIDIA/BigVGAN
A universal neural vocoder that generates high-quality audio waveforms from acoustic features for speech and music synthesis.

Not currently ranked — collecting fresh signals.
star history
BigVGAN is a neural vocoder model published at ICLR 2023, designed to convert acoustic features such as mel-spectrograms into high-fidelity audio waveforms. It uses a GAN-based architecture with custom CUDA kernels for accelerated inference. The model supports speech synthesis, singing voice synthesis, and general audio generation, with pretrained checkpoints available via Hugging Face.
Frequently asked
- What is NVIDIA/BigVGAN?
- A universal neural vocoder that generates high-quality audio waveforms from acoustic features for speech and music synthesis.
- Is BigVGAN open source?
- Yes — NVIDIA/BigVGAN is open source, released under the MIT license.
- What language is BigVGAN written in?
- NVIDIA/BigVGAN is primarily written in Python.
- How popular is BigVGAN?
- NVIDIA/BigVGAN has 1.2k stars on GitHub.
- Where can I find BigVGAN?
- NVIDIA/BigVGAN is on GitHub at https://github.com/NVIDIA/BigVGAN.