Let the data decide: off-policy distillation during continued pre-training analyzes objective-to-capability trade-offs and adaptive objective routing in supervision and performance
Read the original at arxiv.org→arXiv:2607.16246v1 Announce Type: new Abstract: Off-policy distillation is now central to large language model pre-training, yet how training data, objective parameterization, and model capabilities interact remains...
Original headline: "Let the Data Decide: Supervision Analysis, Capability Trade-offs, and Adaptive Objective Routing in Continued Pre-Training via Off-Policy Distillation"