← all repositories

NVIDIA/BigVGAN

A universal neural vocoder that generates high-quality audio waveforms from acoustic features for speech and music synthesis.

1.2k stars Python Image · Video · Audio
BigVGAN
Not currently ranked — collecting fresh signals.
star history

BigVGAN is a neural vocoder model published at ICLR 2023, designed to convert acoustic features such as mel-spectrograms into high-fidelity audio waveforms. It uses a GAN-based architecture with custom CUDA kernels for accelerated inference. The model supports speech synthesis, singing voice synthesis, and general audio generation, with pretrained checkpoints available via Hugging Face.

Frequently asked

What is NVIDIA/BigVGAN?
A universal neural vocoder that generates high-quality audio waveforms from acoustic features for speech and music synthesis.
Is BigVGAN open source?
Yes — NVIDIA/BigVGAN is open source, released under the MIT license.
What language is BigVGAN written in?
NVIDIA/BigVGAN is primarily written in Python.
How popular is BigVGAN?
NVIDIA/BigVGAN has 1.2k stars on GitHub.
Where can I find BigVGAN?
NVIDIA/BigVGAN is on GitHub at https://github.com/NVIDIA/BigVGAN.

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.