NVlabs/Long-RL
A full-stack framework for scaling reinforcement learning training of vision-language models to long video reasoning.

Long-RL addresses the challenge of applying reinforcement learning to long video reasoning in vision-language models. It provides a 104K-sample dataset called LongVideo-Reason with high-quality reasoning annotations across diverse domains, combined with a two-stage training pipeline that handles the computational challenges of long-sequence sequence parallelism. The work produces the LongVILA-R1-7B model, demonstrating effective RL scaling to extended multi-modal contexts.
Frequently asked
- What is NVlabs/Long-RL?
- A full-stack framework for scaling reinforcement learning training of vision-language models to long video reasoning.
- Is Long-RL open source?
- Yes — NVlabs/Long-RL is open source, released under the Apache-2.0 license.
- What language is Long-RL written in?
- NVlabs/Long-RL is primarily written in Python.
- How popular is Long-RL?
- NVlabs/Long-RL has 726 stars on GitHub.
- Where can I find Long-RL?
- NVlabs/Long-RL is on GitHub at https://github.com/NVlabs/Long-RL.