Reasoning fine-tuning induces persistent latent policy states
Read the original at arxiv.org→arXiv:2607.18532v1 Announce Type: new Abstract: Reasoning-specialized language models show large performance gains over base models, yet the internal changes responsible for improved multi-step reasoning remain...
Original headline: "Reasoning Fine-Tuning Induces Persistent Latent Policy States"