gpt-omni/mini-omni2
An open-source multimodal language model enabling real-time voice conversations with image, audio, and text understanding.

Not currently ranked — collecting fresh signals.
star history
Mini-Omni2 is a foundation model designed to replicate GPT-4o-style omni capabilities. It accepts image, audio, and text inputs and produces end-to-end speech-to-speech responses without requiring separate ASR or TTS models. The project includes model weights, inference code, and chat demo functionality for real-time conversational interaction with interruption handling.
Frequently asked
- What is gpt-omni/mini-omni2?
- An open-source multimodal language model enabling real-time voice conversations with image, audio, and text understanding.
- Is mini-omni2 open source?
- Yes — gpt-omni/mini-omni2 is open source, released under the MIT license.
- What language is mini-omni2 written in?
- gpt-omni/mini-omni2 is primarily written in Python.
- How popular is mini-omni2?
- gpt-omni/mini-omni2 has 1.9k stars on GitHub.
- Where can I find mini-omni2?
- gpt-omni/mini-omni2 is on GitHub at https://github.com/gpt-omni/mini-omni2.