StalePO: anchored token-level preference optimization using legacy post-edits in machine translation
Read the original at arxiv.org→arXiv:2609.16340v1 Announce Type: new Abstract: Machine translation systems are periodically upgraded to stronger models, but the available preference signal is human post-edits of an older system's outputs, which...
Original headline: "StalePO: Anchored Token-Level Preference Optimization using Legacy Post-Edits in Machine Translation"
Coverage timeline
- Sep 16, 04:00 UTC arXiv cs.CL lead source StalePO: Anchored Token-Level Preference Optimization using Legacy Post-Edits in Machine Translation