Tail-Likelihood Reinforcement Learning; researchers propose optimizing coverage to distinguish policies with identical mean rewards but different rare high-reward rollout probabilities.
Read the original at arxiv.org→arXiv:2609.02987v1 Announce Type: new Abstract: Reinforcement learning typically optimizes average reward. For generative policies, the average can hide an important distinction: two policies can achieve the same...
Original headline: "Tail-Likelihood Reinforcement Learning"
Coverage timeline
- Sep 4, 04:00 UTC arXiv cs.LG lead source Tail-Likelihood Reinforcement Learning