Recursive self-improvement via on-policy distillation for reasoning
Read the original at arxiv.org→arXiv:2609.30652v1 Announce Type: new Abstract: On-policy distillation (OPD) trains a student model by having it generate trajectories, then matching its next-token predictions with an external teacher's next-token...
Original headline: "Recursive Self-Improvement via On-Policy Distillation for Reasoning"
Coverage timeline
- Sep 28, 04:00 UTC arXiv cs.CL lead source Recursive Self-Improvement via On-Policy Distillation for Reasoning