Rebalancing token importance in language models with TF-IDF weighted cross-entropy loss
Read the original at arxiv.org→arXiv:2609.11029v1 Announce Type: new Abstract: Large language models are typically trained under uniform token weighting, which allows frequent and low-information tokens to dominate learning and can increase the...
Original headline: "Rebalancing Token Importance in Language Models with TF-IDF Weighted Cross-Entropy Loss"
Coverage timeline
- Sep 11, 04:00 UTC arXiv cs.CL lead source Rebalancing Token Importance in Language Models with TF-IDF Weighted Cross-Entropy Loss