QwenLM/Qwen3-Omni
A natively end-to-end multilingual omni-modal foundation model that processes text, audio, images, and video while generating real-time text and speech responses.

Not currently ranked — collecting fresh signals.
star history
Qwen3-Omni is a foundation model that handles multiple modalities in a unified architecture. It processes text, images, audio, and video as inputs and generates both text and natural speech as outputs in real time. The model represents an end-to-end approach to multimodal understanding and generation, released with model weights, demos, and cookbooks by Alibaba Cloud’s Qwen team.
Frequently asked
- What is QwenLM/Qwen3-Omni?
- A natively end-to-end multilingual omni-modal foundation model that processes text, audio, images, and video while generating real-time text and speech responses.
- Is Qwen3-Omni open source?
- Yes — QwenLM/Qwen3-Omni is open source, released under the Apache-2.0 license.
- What language is Qwen3-Omni written in?
- QwenLM/Qwen3-Omni is primarily written in Jupyter Notebook.
- How popular is Qwen3-Omni?
- QwenLM/Qwen3-Omni has 3.9k stars on GitHub.
- Where can I find Qwen3-Omni?
- QwenLM/Qwen3-Omni is on GitHub at https://github.com/QwenLM/Qwen3-Omni.