Liuziyu77/Visual-RFT
Visual-RFT applies reinforcement learning fine-tuning to vision-language models using a GRPO-based framework with rule-based verifiable rewards.

Not currently ranked — collecting fresh signals.
star history
The repository provides the first comprehensive adaptation of Deepseek-R1’s reinforcement learning strategy to the multimodal domain. It fine-tunes Qwen2-VL-2/7B models through a GRPO-based framework with rule-based verifiable rewards, enhancing performance across various visual perception tasks including open vocabulary detection, few-shot detection, and reasoning grounding.
Frequently asked
- What is Liuziyu77/Visual-RFT?
- Visual-RFT applies reinforcement learning fine-tuning to vision-language models using a GRPO-based framework with rule-based verifiable rewards.
- Is Visual-RFT open source?
- Yes — Liuziyu77/Visual-RFT is open source, released under the Apache-2.0 license.
- What language is Visual-RFT written in?
- Liuziyu77/Visual-RFT is primarily written in Jupyter Notebook.
- How popular is Visual-RFT?
- Liuziyu77/Visual-RFT has 2.3k stars on GitHub.
- Where can I find Visual-RFT?
- Liuziyu77/Visual-RFT is on GitHub at https://github.com/Liuziyu77/Visual-RFT.