← all repositories

uclaml/SPPO

Self-Play Preference Optimization is a self-play framework for language model alignment with a new learning objective, released with trained model weights.

590 stars Python Language ModelsML Frameworks
SPPO
Not currently ranked — collecting fresh signals.
star history

The repository provides the official implementation of SPPO, a self-play-based method for aligning language models using a novel learning objective derived from game theory. It includes training scripts for fine-tuning LLMs, evaluation pipelines on benchmarks like AlpacaEval 2.0 and Open LLM Leaderboard, and released model checkpoints. The approach frames alignment as a competitive two-player game where the model improves by playing against itself.

Frequently asked

What is uclaml/SPPO?
Self-Play Preference Optimization is a self-play framework for language model alignment with a new learning objective, released with trained model weights.
Is SPPO open source?
Yes — uclaml/SPPO is open source, released under the Apache-2.0 license.
What language is SPPO written in?
uclaml/SPPO is primarily written in Python.
How popular is SPPO?
uclaml/SPPO has 590 stars on GitHub.
Where can I find SPPO?
uclaml/SPPO is on GitHub at https://github.com/uclaml/SPPO.

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.