← all repositories

FunAudioLLM/ThinkSound

ThinkSound is a unified framework for generating audio from any modality (text, video, images) guided by Chain-of-Thought reasoning using PyTorch and flow matching.

1.4k stars Python Image · Video · Audio
ThinkSound
Not currently ranked — collecting fresh signals.
star history

ThinkSound presents a unified approach to any-modality-to-audio generation, leveraging Chain-of-Thought reasoning to guide the generation process. The framework uses flow matching as its core synthesis mechanism and supports diverse inputs including text descriptions, video, and images. This NeurIPS 2025 publication from FunAudioLLM targets tasks like foley sound synthesis, sound effect generation, and multimodal audio creation.

Frequently asked

What is FunAudioLLM/ThinkSound?
ThinkSound is a unified framework for generating audio from any modality (text, video, images) guided by Chain-of-Thought reasoning using PyTorch and flow matching.
Is ThinkSound open source?
Yes — FunAudioLLM/ThinkSound is an open-source project tracked on heatdrop.
What language is ThinkSound written in?
FunAudioLLM/ThinkSound is primarily written in Python.
How popular is ThinkSound?
FunAudioLLM/ThinkSound has 1.4k stars on GitHub.
Where can I find ThinkSound?
FunAudioLLM/ThinkSound is on GitHub at https://github.com/FunAudioLLM/ThinkSound.

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.