← all repositories

QwenLM/Qwen3-Omni

A natively end-to-end multilingual omni-modal foundation model that processes text, audio, images, and video while generating real-time text and speech responses.

3.9k stars Jupyter Notebook Language ModelsImage · Video · Audio
Qwen3-Omni
Not currently ranked — collecting fresh signals.
star history

Qwen3-Omni is a foundation model that handles multiple modalities in a unified architecture. It processes text, images, audio, and video as inputs and generates both text and natural speech as outputs in real time. The model represents an end-to-end approach to multimodal understanding and generation, released with model weights, demos, and cookbooks by Alibaba Cloud’s Qwen team.

Frequently asked

What is QwenLM/Qwen3-Omni?
A natively end-to-end multilingual omni-modal foundation model that processes text, audio, images, and video while generating real-time text and speech responses.
Is Qwen3-Omni open source?
Yes — QwenLM/Qwen3-Omni is open source, released under the Apache-2.0 license.
What language is Qwen3-Omni written in?
QwenLM/Qwen3-Omni is primarily written in Jupyter Notebook.
How popular is Qwen3-Omni?
QwenLM/Qwen3-Omni has 3.9k stars on GitHub.
Where can I find Qwen3-Omni?
QwenLM/Qwen3-Omni is on GitHub at https://github.com/QwenLM/Qwen3-Omni.

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.