gpt-omni/mini-omni
An open-source multimodal LLM that enables real-time speech-to-speech conversation with streaming audio output and concurrent text/audio generation.

Not currently ranked — collecting fresh signals.
star history
Mini-Omni is a multimodal large language model designed for real-time voice conversation. It processes speech input directly without requiring separate ASR or TTS models, enabling true end-to-end speech-to-speech interaction. The model can generate text and audio simultaneously while thinking, and supports streaming audio output for natural conversational experiences.
Frequently asked
- What is gpt-omni/mini-omni?
- An open-source multimodal LLM that enables real-time speech-to-speech conversation with streaming audio output and concurrent text/audio generation.
- Is mini-omni open source?
- Yes — gpt-omni/mini-omni is open source, released under the MIT license.
- What language is mini-omni written in?
- gpt-omni/mini-omni is primarily written in Python.
- How popular is mini-omni?
- gpt-omni/mini-omni has 3.6k stars on GitHub.
- Where can I find mini-omni?
- gpt-omni/mini-omni is on GitHub at https://github.com/gpt-omni/mini-omni.