Paper claims RL for reasoning changes only 1-3% of tokens and replicates gains without RL at ~1000x less compute
Read the original at old.reddit.com→submitted by /u/juanviera23 [link] [comments]
Original headline: "Paper claims RL for reasoning only changes 1-3% of tokens, and they replicate the gains without RL at ~1000x less compute"
Coverage timeline
- Aug 16, 11:21 UTC r/LocalLLaMA lead source Paper claims RL for reasoning only changes 1-3% of tokens, and they replicate the gains without RL at ~1000x less compute