Weak-to-Strong on-policy distillation improves alignment when no larger teacher exists
Read the original at arxiv.org→arXiv:2607.26246v1 Announce Type: new Abstract: On-policy distillation (OPD), which aligns a student with the teacher's token-level distribution on the student's own rollouts, is an effective paradigm for...
Original headline: "Weak-to-Strong On-Policy Distillation"
Coverage timeline
- Jul 30, 04:00 UTC arXiv cs.LG lead source Weak-to-Strong On-Policy Distillation