Regularized emphatic temporal-difference learning shows stability under constant stepsizes
Read the original at arxiv.org→arXiv:2609.19170v1 Announce Type: new Abstract: Emphatic temporal-difference learning (ETD) stabilizes the expected off-policy TD update and changes its projection geometry, but neither property determines...
Original headline: "Regularized Emphatic Temporal-Difference Learning: Stability under Constant Stepsizes"
Coverage timeline
- Sep 18, 04:00 UTC arXiv cs.AI lead source Regularized Emphatic Temporal-Difference Learning: Stability under Constant Stepsizes