Normalized Rewards for Preference Optimization
Read the original at arxiv.org→arXiv:2607.16240v1 Announce Type: new Abstract: Direct Alignment Algorithms (DAAs) such as DPO have become a common way to post-train and align LLMs with human preferences. However, DAAs have been observed to...
Coverage timeline
- Jul 21, 04:00 UTC arXiv cs.LG lead source Normalized Rewards for Preference Optimization
- Jul 22, 04:00 UTC arXiv cs.LG TD-DPO: Difference-Aware Preference Optimization for Mitigating Sycophancy in Clinical Autism Intervention Dialogue