Executable walkthrough induction from sparse-reward trajectories enables learning state-conditioned procedures for long-horizon agents
Read the original at arxiv.org→arXiv:2609.22120v1 Announce Type: new Abstract: Test-time self-evolving agents improve by reusing past experience, yet sparse-reward trajectories contain failures, loops, and detours, while summaries often omit the...
Original headline: "Success Leaves Detours: Learning Executable Walkthroughs for Long-Horizon Agents"
Coverage timeline
- Sep 22, 04:00 UTC arXiv cs.LG lead source Success Leaves Detours: Learning Executable Walkthroughs for Long-Horizon Agents