tatsu-lab/alpaca_farm
A simulation framework for developing and benchmarking RLHF methods without collecting real human feedback data.

Not currently ranked — collecting fresh signals.
star history
AlpacaFarm provides a low-cost environment for research on learning from human feedback, specifically for instruction-following and alignment of language models. It simulates the RLHF pipeline including preference annotation, allowing researchers to develop and evaluate RLHF methods using automated annotators like GPT-4. The project includes reference implementations of PPO, DPO, and other training algorithms.
Frequently asked
- What is tatsu-lab/alpaca_farm?
- A simulation framework for developing and benchmarking RLHF methods without collecting real human feedback data.
- Is alpaca_farm open source?
- Yes — tatsu-lab/alpaca_farm is open source, released under the Apache-2.0 license.
- What language is alpaca_farm written in?
- tatsu-lab/alpaca_farm is primarily written in Python.
- How popular is alpaca_farm?
- tatsu-lab/alpaca_farm has 845 stars on GitHub.
- Where can I find alpaca_farm?
- tatsu-lab/alpaca_farm is on GitHub at https://github.com/tatsu-lab/alpaca_farm.