It Takes 8 tokens: weak-to-strong off-policy RL via auxiliary branches
Read the original at arxiv.org→arXiv:2607.16205v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards has emerged as a standard approach for enhancing reasoning in large language models, which typically optimizes the...
Original headline: "It Takes 8 Tokens: Weak-to-Strong Off-Policy RL via Auxiliary Branches"