Differentially private inference-time alignment via calibrated noise to reward scores
Read the original at arxiv.org→arXiv:2608.26324v1 Announce Type: new Abstract: Best-of-N (BoN) sampling is the simplest and most widely deployed inference-time alignment strategy, but it suffers from two distinct problems: reward hacking, in...
Original headline: "Privacy Without Regret: Differentially Private Inference-Time Alignment"
Coverage timeline
- Aug 28, 04:00 UTC arXiv cs.LG lead source Privacy Without Regret: Differentially Private Inference-Time Alignment