Rater state bias in RLHF preference data; an audit framework
Read the original at arxiv.org→arXiv:2607.16195v1 Announce Type: new Abstract: We identify a structured confound in Reinforcement Learning from Human Feedback (RLHF). Pairwise preference labels are intended to reflect the compared outputs, but...
Original headline: "Rater State Bias in RLHF Preference Data: An Audit Framework"