Arrive and Survive: Scaling Safe Goal-Conditioned Policy Learning from One-Bit Failure Signals
Read the original at arxiv.org→arXiv:2608.26571v1 Announce Type: new Abstract: Contrastive reinforcement learning (CRL) scales effectively in goal-conditioned tasks by casting policy learning into a self-supervised contrastive objective. However,...
Coverage timeline
- Aug 28, 04:00 UTC arXiv cs.LG lead source Arrive and Survive: Scaling Safe Goal-Conditioned Policy Learning from One-Bit Failure Signals