Which tokens should SFT actually learn? A token-trimming perspective on mathematical reasoning
Read the original at arxiv.org→arXiv:2609.09707v1 Announce Type: new Abstract: Supervised fine-tuning (SFT) applies a uniform cross-entropy loss to all target tokens, even though different tokens provide unequal learning signals for mathematical...
Original headline: "Which Tokens Should SFT Actually Learn? A Token-Trimming Perspective on Mathematical Reasoning"
Coverage timeline
- Sep 10, 04:00 UTC arXiv cs.AI lead source Which Tokens Should SFT Actually Learn? A Token-Trimming Perspective on Mathematical Reasoning