Near-optimal sample complexity for recursive entropic risk RL with a generative model
Read the original at arxiv.org→arXiv:2610.06931v1 Announce Type: new Abstract: In this paper, we study the sample complexities of value and policy learning in finite discounted Markov decision processes (MDPs) under recursive entropic risk...
Original headline: "Near-Optimal Sample Complexity for Recursive Entropic Risk Reinforcement Learning with a Generative Model"
Coverage timeline
- Oct 7, 04:00 UTC arXiv cs.LG lead source Near-Optimal Sample Complexity for Recursive Entropic Risk Reinforcement Learning with a Generative Model