Fork where the model changes its mind: belief-shift branching for tree-structured reinforcement learning
Read the original at arxiv.org→arXiv:2609.11061v1 Announce Type: new Abstract: Tree-structured rollouts give critic-free reinforcement learning with verifiable rewards (RLVR) step-level credit: fork a chain at an intermediate point, and sibling...
Original headline: "Fork Where the Model Changes Its Mind: Belief-Shift Branching for Tree-Structured Reinforcement Learning"
Coverage timeline
- Sep 12, 04:00 UTC arXiv cs.AI lead source Fork Where the Model Changes Its Mind: Belief-Shift Branching for Tree-Structured Reinforcement Learning