Finite constant frontiers and auditable regret certificates for average-reward reinforcement learning
Read the original at arxiv.org→arXiv:2608.07725v1 Announce Type: new Abstract: Average-reward reinforcement-learning regret is known up to logarithmic factors, but the numerical content of published guarantees is difficult to compare because...
Original headline: "Finite Constant Frontiers and Auditable Regret Certificates for Average-Reward Reinforcement Learning"
Coverage timeline
- Aug 11, 04:00 UTC arXiv cs.LG lead source Finite Constant Frontiers and Auditable Regret Certificates for Average-Reward Reinforcement Learning