turningpoint-ai/VisualThinker-R1-Zero
Reinforcement learning post-training for visual reasoning that replicates DeepSeek-R1-Zero's emergent reasoning on a 2B multimodal model.

VisualThinker-R1-Zero applies GRPO-based reinforcement learning to train Qwen2-VL-2B on visual reasoning tasks without supervised fine-tuning or reward models. The project demonstrates emergent self-reflection and correction behaviors in visual reasoning, successfully reproducing the ‘aha moment’ and increasing response length observed in DeepSeek-R1-Zero. This enables reasoning capabilities to emerge from pure RL training on vision-centric tasks.
Frequently asked
- What is turningpoint-ai/VisualThinker-R1-Zero?
- Reinforcement learning post-training for visual reasoning that replicates DeepSeek-R1-Zero's emergent reasoning on a 2B multimodal model.
- Is VisualThinker-R1-Zero open source?
- Yes — turningpoint-ai/VisualThinker-R1-Zero is an open-source project tracked on heatdrop.
- What language is VisualThinker-R1-Zero written in?
- turningpoint-ai/VisualThinker-R1-Zero is primarily written in Python.
- How popular is VisualThinker-R1-Zero?
- turningpoint-ai/VisualThinker-R1-Zero has 624 stars on GitHub.
- Where can I find VisualThinker-R1-Zero?
- turningpoint-ai/VisualThinker-R1-Zero is on GitHub at https://github.com/turningpoint-ai/VisualThinker-R1-Zero.