← all repositories

Tencent-Hunyuan/MixGRPO

Tencent Hunyuan's research on improving Group Relative Policy Optimization for flow-matching diffusion models via mixed ODE-SDE sampling.

MixGRPO
Not currently ranked — collecting fresh signals.
star history

MixGRPO is a reinforcement learning training method designed to improve the efficiency of GRPO for flow-based generative models. It introduces a mixed ODE-SDE (Ordinary Differential Equation - Stochastic Differential Equation) sampling strategy to better balance exploration and exploitation during training. The approach targets diffusion models used in generative tasks, aiming to unlock more efficient policy optimization by jointly optimizing sampling trajectories and reward signals.

Frequently asked

What is Tencent-Hunyuan/MixGRPO?
Tencent Hunyuan's research on improving Group Relative Policy Optimization for flow-matching diffusion models via mixed ODE-SDE sampling.
Is MixGRPO open source?
Yes — Tencent-Hunyuan/MixGRPO is an open-source project tracked on heatdrop.
What language is MixGRPO written in?
Tencent-Hunyuan/MixGRPO is primarily written in Python.
How popular is MixGRPO?
Tencent-Hunyuan/MixGRPO has 1.2k stars on GitHub.
Where can I find MixGRPO?
Tencent-Hunyuan/MixGRPO is on GitHub at https://github.com/Tencent-Hunyuan/MixGRPO.

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.