← all repositories

hkchengrex/MMAudio

A deep learning model that generates synchronized audio from video and/or text inputs.

2.2k stars Python Image · Video · Audio
MMAudio
Not currently ranked — collecting fresh signals.
star history

MMAudio is a generative model for video-to-audio synthesis that takes video frames and/or text as input to produce matching audio. It uses multimodal joint training across audio-visual and audio-text datasets to enable high-quality audio generation. A synchronization module aligns the generated audio with the video frames for temporal coherence. The project provides pretrained models and interactive demos via Huggingface, Colab, and Replicate.

Frequently asked

What is hkchengrex/MMAudio?
A deep learning model that generates synchronized audio from video and/or text inputs.
Is MMAudio open source?
Yes — hkchengrex/MMAudio is open source, released under the MIT license.
What language is MMAudio written in?
hkchengrex/MMAudio is primarily written in Python.
How popular is MMAudio?
hkchengrex/MMAudio has 2.2k stars on GitHub.
Where can I find MMAudio?
hkchengrex/MMAudio is on GitHub at https://github.com/hkchengrex/MMAudio.

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.