Reasoning fine-tuning induces persistent latent policy states
Read the original at arxiv.org→arXiv:2607.18532v1 Announce Type: new Abstract: Reasoning-specialized language models show large performance gains over base models, yet the internal changes responsible for improved multi-step reasoning remain...
Original headline: "Reasoning Fine-Tuning Induces Persistent Latent Policy States"
Coverage timeline
- Jul 22, 04:00 UTC arXiv cs.CL lead source Reasoning Fine-Tuning Induces Persistent Latent Policy States
- Jul 22, 04:00 UTC arXiv cs.CL Reasoning Error from Known Fact: Step-Level Self-Consistency Group Relative Policy Optimization for LLM
- Jul 23, 04:00 UTC arXiv cs.CL SLPO: Scaling Latent Reasoning via a Surrogate Policy