← all repositories

Joyce94/LLM-RLHF-Tuning

A complete RLHF training framework implementing SFT, reward modeling, PPO, and DPO for fine-tuning language models with LoRA/PEFT.

LLM-RLHF-Tuning
Not currently ranked — collecting fresh signals.
star history

This repository provides a from-scratch implementation of the three-stage RLHF (Reinforcement Learning from Human Feedback) training pipeline for large language models. It supports supervised fine-tuning (SFT), reward model (RM) training, PPO (Proximal Policy Optimization) training, and DPO (Direct Preference Optimization) training. The framework leverages PEFT and LoRA for parameter-efficient fine-tuning, supporting LLaMA, LLaMA2, and Alpaca models with distributed training via accelerate and DeepSpeed.

Frequently asked

What is Joyce94/LLM-RLHF-Tuning?
A complete RLHF training framework implementing SFT, reward modeling, PPO, and DPO for fine-tuning language models with LoRA/PEFT.
Is LLM-RLHF-Tuning open source?
Yes — Joyce94/LLM-RLHF-Tuning is an open-source project tracked on heatdrop.
What language is LLM-RLHF-Tuning written in?
Joyce94/LLM-RLHF-Tuning is primarily written in Python.
How popular is LLM-RLHF-Tuning?
Joyce94/LLM-RLHF-Tuning has 451 stars on GitHub.
Where can I find LLM-RLHF-Tuning?
Joyce94/LLM-RLHF-Tuning is on GitHub at https://github.com/Joyce94/LLM-RLHF-Tuning.

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.