inclusionAI/UI-Venus
A native UI agent for precise GUI element grounding and navigation using screenshot-based multimodal LLMs.

UI-Venus is a unified, end-to-end GUI agent that performs precise GUI element grounding and effective navigation using only screenshots as input. The system includes 2B/8B dense and 30B-A3B MoE model variants trained with mid-stage training on 10B tokens across 30+ datasets and online reinforcement learning for long-horizon navigation. It achieves state-of-the-art performance on benchmarks including ScreenSpot-Pro (69.6%), VenusBench-GD (75.0%), and AndroidWorld (77.6%), with demonstrated navigation across 40+ Chinese mobile applications.
Frequently asked
- What is inclusionAI/UI-Venus?
- A native UI agent for precise GUI element grounding and navigation using screenshot-based multimodal LLMs.
- Is UI-Venus open source?
- Yes — inclusionAI/UI-Venus is an open-source project tracked on heatdrop.
- What language is UI-Venus written in?
- inclusionAI/UI-Venus is primarily written in Python.
- How popular is UI-Venus?
- inclusionAI/UI-Venus has 1k stars on GitHub.
- Where can I find UI-Venus?
- inclusionAI/UI-Venus is on GitHub at https://github.com/inclusionAI/UI-Venus.