Policy learning with mu-resets: the sample complexity under policy realizability for the Kakade–Langford interaction protocol
Read the original at arxiv.org→arXiv:2608.07772v1 Announce Type: new Abstract: We study policy-based reinforcement learning under the $\mu$-resets interaction protocol of Kakade and Langford [KL02]. This interaction protocol enables the learner...
Original headline: "The Sample Complexity of Policy Learning with Mu-Resets"
Coverage timeline
- Aug 11, 04:00 UTC arXiv cs.LG lead source The Sample Complexity of Policy Learning with Mu-Resets