CODS: Iterative Bellman-residual data selection for reusable offline reinforcement learning
Read the original at arxiv.org→arXiv:2608.07719v1 Announce Type: new Abstract: Offline reinforcement learning repeatedly trains policies from a fixed transition pool, making redundant data costly across seeds and hyperparameters, while naive...
Original headline: "CODS: Iterative Bellman-Residual Data Selection for Reusable Offline Reinforcement Learning"
Coverage timeline
- Aug 11, 04:00 UTC arXiv cs.LG lead source CODS: Iterative Bellman-Residual Data Selection for Reusable Offline Reinforcement Learning