Contextual Early Reward predicts terminal reward from trajectory prefixes for long-horizon coding agents
Read the original at arxiv.org→arXiv:2609.31995v1 Announce Type: new Abstract: Long-horizon coding agents receive verifiable rewards only after completing expensive sequences of tool calls. This increases inference cost, amplifies early wrong...
Original headline: "Before the Rollout Ends: Early Terminal Reward Prediction for Long-horizon Coding Agents"
Coverage timeline
- Sep 29, 04:00 UTC arXiv cs.CL lead source Before the Rollout Ends: Early Terminal Reward Prediction for Long-horizon Coding Agents