From weak data to strong policy: Q-Targets enable provable in-context reinforcement learning
Read the original at arxiv.org→arXiv:2609.30391v1 Announce Type: new Abstract: Existing in-context reinforcement learning methods mainly pretrain Transformers with supervised behavior-prediction objectives. This enables task inference from...
Original headline: "From Weak Data to Strong Policy: Q-Targets Enable Provable In-Context Reinforcement Learning"
Coverage timeline
- Sep 28, 04:00 UTC arXiv cs.LG lead source From Weak Data to Strong Policy: Q-Targets Enable Provable In-Context Reinforcement Learning