Training and evaluating ethical reinforcement learning agents on per-episode distributions
Read the original at arxiv.org→arXiv:2608.14642v1 Announce Type: new Abstract: Reinforcement Learning (RL) agents trained on a single reward signal exploit the gap between the designed reward and the intended behavior. This is particularly a...
Original headline: "Training and Evaluating Ethical Reinforcement Learning Agents on Per-Episode Distributions"
Coverage timeline
- Aug 18, 04:00 UTC arXiv cs.LG lead source Training and Evaluating Ethical Reinforcement Learning Agents on Per-Episode Distributions