LC-GRPO: Bridging train-inference gap for flow-based GRPO with Langevin correction
Read the original at arxiv.org→arXiv:2608.05600v1 Announce Type: new Abstract: Flow-based generative models are typically sampled by solving a deterministic ordinary differential equation (ODE), whereas online reinforcement learning requires...
Original headline: "LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction"
Coverage timeline
- Aug 7, 04:00 UTC arXiv cs.LG lead source LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction