Pretraining and midtraining enable learning from rewards, by providing information and computation for reward adaptation; understanding through sequential state computation and contextual memory.
Read the original at arxiv.org→arXiv:2609.38446v1 Announce Type: new Abstract: A reward can identify a correct answer while leaving the computation needed for new inputs undetermined. We study how pretraining and midtraining supply the...
Original headline: "What Pretraining and Midtraining Make Learnable from Rewards?"
Coverage timeline
- Oct 1, 04:00 UTC arXiv cs.LG lead source What Pretraining and Midtraining Make Learnable from Rewards?