← all repositories

gpt-omni/mini-omni

An open-source multimodal LLM that enables real-time speech-to-speech conversation with streaming audio output and concurrent text/audio generation.

mini-omni
Not currently ranked — collecting fresh signals.
star history

Mini-Omni is a multimodal large language model designed for real-time voice conversation. It processes speech input directly without requiring separate ASR or TTS models, enabling true end-to-end speech-to-speech interaction. The model can generate text and audio simultaneously while thinking, and supports streaming audio output for natural conversational experiences.

Frequently asked

What is gpt-omni/mini-omni?
An open-source multimodal LLM that enables real-time speech-to-speech conversation with streaming audio output and concurrent text/audio generation.
Is mini-omni open source?
Yes — gpt-omni/mini-omni is open source, released under the MIT license.
What language is mini-omni written in?
gpt-omni/mini-omni is primarily written in Python.
How popular is mini-omni?
gpt-omni/mini-omni has 3.6k stars on GitHub.
Where can I find mini-omni?
gpt-omni/mini-omni is on GitHub at https://github.com/gpt-omni/mini-omni.

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.