Abstention as an action can kill both the reward gradient and the KL anchor: collapse law and repair for error-penalized reinforcement learning
Read the original at arxiv.org→arXiv:2608.00301v1 Announce Type: new Abstract: Error-penalized scoring rules ($+1$ for a correct answer, $-\lambda$ for a wrong one, $0$ for abstaining) are increasingly prescribed against hallucination: a rational...
Original headline: "Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning"
Coverage timeline
- Aug 4, 04:00 UTC arXiv cs.LG lead source Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning