MoonshotAI/Kimi-VL
Open-source Mixture-of-Experts vision-language model with 128K context window and autonomous agent capabilities.

Kimi-VL is an efficient open-source vision-language model (VLM) using Mixture-of-Experts architecture in its language decoder, activating only 2.8B parameters. It supports multimodal reasoning, long-context understanding up to 128K tokens, and demonstrates strong agent capabilities in multi-turn interactions such as OSWorld. The model includes a native-resolution vision encoder (MoonViT) and achieves competitive performance against GPT-4o-mini, Qwen2.5-VL-7B, and other efficient VLMs.
Frequently asked
- What is MoonshotAI/Kimi-VL?
- Open-source Mixture-of-Experts vision-language model with 128K context window and autonomous agent capabilities.
- Is Kimi-VL open source?
- Yes — MoonshotAI/Kimi-VL is open source, released under the MIT license.
- How popular is Kimi-VL?
- MoonshotAI/Kimi-VL has 1.2k stars on GitHub.
- Where can I find Kimi-VL?
- MoonshotAI/Kimi-VL is on GitHub at https://github.com/MoonshotAI/Kimi-VL.