NVIDIA/audio-flamingo
NVIDIA's Audio Flamingo is a series of open-source multimodal LLMs that understand speech and music through natural language interactions.

Not currently ranked — collecting fresh signals.
star history
Audio Flamingo provides PyTorch implementations of large language models trained to understand audio through text queries. The models support audio captioning, question answering, reasoning, and long-audio understanding across speech and music domains. Multiple versions have been published at top ML venues (ICML, NeurIPS), with the latest being fully open-sourced for research use.
Frequently asked
- What is NVIDIA/audio-flamingo?
- NVIDIA's Audio Flamingo is a series of open-source multimodal LLMs that understand speech and music through natural language interactions.
- Is audio-flamingo open source?
- Yes — NVIDIA/audio-flamingo is an open-source project tracked on heatdrop.
- How popular is audio-flamingo?
- NVIDIA/audio-flamingo has 1.1k stars on GitHub.
- Where can I find audio-flamingo?
- NVIDIA/audio-flamingo is on GitHub at https://github.com/NVIDIA/audio-flamingo.