← all repositories

RLHFlow/RLHF-Reward-Modeling

A collection of training recipes for reward models used in RLHF-based LLM alignment, including Bradley-Terry, pairwise, multi-objective, and process-supervised approaches.

RLHF-Reward-Modeling
Not currently ranked — collecting fresh signals.
star history

The repository provides implementations of various reward modeling techniques for training LLMs via Reinforcement Learning from Human Feedback (RLHF). It includes classic Bradley-Terry reward modeling, pairwise preference models that predict response preference from prompt-response pairs, multi-objective reward models using mixture-of-experts aggregation, and process-supervised reward models for mathematical reasoning. Each approach includes code, data, hyperparameters, and references to associated research papers.

Frequently asked

What is RLHFlow/RLHF-Reward-Modeling?
A collection of training recipes for reward models used in RLHF-based LLM alignment, including Bradley-Terry, pairwise, multi-objective, and process-supervised approaches.
Is RLHF-Reward-Modeling open source?
Yes — RLHFlow/RLHF-Reward-Modeling is open source, released under the Apache-2.0 license.
What language is RLHF-Reward-Modeling written in?
RLHFlow/RLHF-Reward-Modeling is primarily written in Python.
How popular is RLHF-Reward-Modeling?
RLHFlow/RLHF-Reward-Modeling has 1.5k stars on GitHub.
Where can I find RLHF-Reward-Modeling?
RLHFlow/RLHF-Reward-Modeling is on GitHub at https://github.com/RLHFlow/RLHF-Reward-Modeling.

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.