Simple-OPD: Demystifying warm-up for on-policy distillation
Read the original at arxiv.org→arXiv:2608.06802v1 Announce Type: new Abstract: On-policy distillation (OPD) trains a student on its own rollouts with token-level supervision from teacher models, but its effectiveness can depend strongly on the...
Original headline: "Simple-OPD: Demystifying Warm-up for On-policy Distillation"
Coverage timeline
- Aug 10, 04:00 UTC arXiv cs.CL lead source Simple-OPD: Demystifying Warm-up for On-policy Distillation