hkchengrex/MMAudio
A deep learning model that generates synchronized audio from video and/or text inputs.

Not currently ranked — collecting fresh signals.
star history
MMAudio is a generative model for video-to-audio synthesis that takes video frames and/or text as input to produce matching audio. It uses multimodal joint training across audio-visual and audio-text datasets to enable high-quality audio generation. A synchronization module aligns the generated audio with the video frames for temporal coherence. The project provides pretrained models and interactive demos via Huggingface, Colab, and Replicate.
Frequently asked
- What is hkchengrex/MMAudio?
- A deep learning model that generates synchronized audio from video and/or text inputs.
- Is MMAudio open source?
- Yes — hkchengrex/MMAudio is open source, released under the MIT license.
- What language is MMAudio written in?
- hkchengrex/MMAudio is primarily written in Python.
- How popular is MMAudio?
- hkchengrex/MMAudio has 2.2k stars on GitHub.
- Where can I find MMAudio?
- hkchengrex/MMAudio is on GitHub at https://github.com/hkchengrex/MMAudio.