← all repositories

Text-to-Audio/Make-An-Audio

A conditional diffusion probabilistic model that generates high-fidelity audio from text and other modality inputs.

669 stars Python Image · Video · Audio
Make-An-Audio
Not currently ranked — collecting fresh signals.
star history

Make-An-Audio is a generative AI system that produces audio from text prompts using latent diffusion models. It employs a prompt-enhanced diffusion approach for conditioning and supports generation from text and video modalities. The repository provides a PyTorch implementation along with pretrained models for audio generation and audio inpainting tasks.

Frequently asked

What is Text-to-Audio/Make-An-Audio?
A conditional diffusion probabilistic model that generates high-fidelity audio from text and other modality inputs.
Is Make-An-Audio open source?
Yes — Text-to-Audio/Make-An-Audio is open source, released under the MIT license.
What language is Make-An-Audio written in?
Text-to-Audio/Make-An-Audio is primarily written in Python.
How popular is Make-An-Audio?
Text-to-Audio/Make-An-Audio has 669 stars on GitHub.
Where can I find Make-An-Audio?
Text-to-Audio/Make-An-Audio is on GitHub at https://github.com/Text-to-Audio/Make-An-Audio.

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.