← all repositories

gpt-omni/mini-omni2

An open-source multimodal language model enabling real-time voice conversations with image, audio, and text understanding.

mini-omni2
Not currently ranked — collecting fresh signals.
star history

Mini-Omni2 is a foundation model designed to replicate GPT-4o-style omni capabilities. It accepts image, audio, and text inputs and produces end-to-end speech-to-speech responses without requiring separate ASR or TTS models. The project includes model weights, inference code, and chat demo functionality for real-time conversational interaction with interruption handling.

Frequently asked

What is gpt-omni/mini-omni2?
An open-source multimodal language model enabling real-time voice conversations with image, audio, and text understanding.
Is mini-omni2 open source?
Yes — gpt-omni/mini-omni2 is open source, released under the MIT license.
What language is mini-omni2 written in?
gpt-omni/mini-omni2 is primarily written in Python.
How popular is mini-omni2?
gpt-omni/mini-omni2 has 1.9k stars on GitHub.
Where can I find mini-omni2?
gpt-omni/mini-omni2 is on GitHub at https://github.com/gpt-omni/mini-omni2.

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.