A shared learning rate is not a neutral control in selective on-policy distillation.
Read the original at arxiv.org→arXiv:2609.22109v1 Announce Type: new Abstract: Selective on-policy distillation trains a student only at the token positions a selector scores highest, and the literature compares selectors under a single shared...
Original headline: "A Shared Learning Rate Is Not a Neutral Control in Selective On-Policy Distillation"
Coverage timeline
- Sep 22, 04:00 UTC arXiv cs.LG lead source A Shared Learning Rate Is Not a Neutral Control in Selective On-Policy Distillation