PRIME-RL/PRIME
A scalable reinforcement learning framework for training language models to reason more effectively using implicit process rewards.

Not currently ranked — collecting fresh signals.
star history
PRIME implements process reinforcement learning to improve LLM reasoning by generating implicit reward signals during multi-step reasoning tasks. The approach focuses on scalable RL training for language models, integrated with frameworks like veRL. It includes training recipes, evaluation benchmarks, and model weights on Hugging Face for reproducing the method.
Frequently asked
- What is PRIME-RL/PRIME?
- A scalable reinforcement learning framework for training language models to reason more effectively using implicit process rewards.
- Is PRIME open source?
- Yes — PRIME-RL/PRIME is open source, released under the Apache-2.0 license.
- What language is PRIME written in?
- PRIME-RL/PRIME is primarily written in Python.
- How popular is PRIME?
- PRIME-RL/PRIME has 1.9k stars on GitHub.
- Where can I find PRIME?
- PRIME-RL/PRIME is on GitHub at https://github.com/PRIME-RL/PRIME.