Preference coverage collapse from hindsight relabeling in multi-objective reinforcement learning
Read the original at arxiv.org→arXiv:2609.26918v1 Announce Type: new Abstract: Hindsight relabeling which retroactively replacing a transition's goal with the outcome the agent actually achieved is an effective tool for improving...
Original headline: "On Preference Coverage Collapse from Hindsight Relabeling in Multi-Objective Reinforcement Learning"
Coverage timeline
- Sep 24, 04:00 UTC arXiv cs.LG lead source On Preference Coverage Collapse from Hindsight Relabeling in Multi-Objective Reinforcement Learning