← all repositories

chi2liu/ABC-GRPO

ABC-GRPO is a reinforcement learning algorithm variant that introduces four independent clipping boundaries to improve stability and generalization when training LLMs like Qwen3 with GRPO.

442 stars Python ML FrameworksLanguage Models
ABC-GRPO
Not currently ranked — collecting fresh signals.
star history

The project implements Adaptive-Boundary-Clipping GRPO, an asymmetric refinement of the standard GRPO reinforcement learning algorithm for LLM training. It replaces GRPO’s two conditional clipping boundaries with four independent parameters (ε₁, ε₂, ε₃, ε₄) that provide unconditional bounds across all quadrants of the advantage space. The method maintains higher entropy during training to prevent premature convergence, and evaluation on mathematical reasoning tasks with Qwen3 models demonstrates superior performance over standard GRPO.

Frequently asked

What is chi2liu/ABC-GRPO?
ABC-GRPO is a reinforcement learning algorithm variant that introduces four independent clipping boundaries to improve stability and generalization when training LLMs like Qwen3 with GRPO.
Is ABC-GRPO open source?
Yes — chi2liu/ABC-GRPO is open source, released under the Apache-2.0 license.
What language is ABC-GRPO written in?
chi2liu/ABC-GRPO is primarily written in Python.
How popular is ABC-GRPO?
chi2liu/ABC-GRPO has 442 stars on GitHub.
Where can I find ABC-GRPO?
chi2liu/ABC-GRPO is on GitHub at https://github.com/chi2liu/ABC-GRPO.

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.