Rater state bias in RLHF preference data; an audit framework
Read the original at arxiv.org→arXiv:2607.16195v1 Announce Type: new Abstract: We identify a structured confound in Reinforcement Learning from Human Feedback (RLHF). Pairwise preference labels are intended to reflect the compared outputs, but...
Original headline: "Rater State Bias in RLHF Preference Data: An Audit Framework"
Coverage timeline
- Jul 21, 04:00 UTC arXiv cs.AI lead source Rater State Bias in RLHF Preference Data: An Audit Framework
- Jul 23, 04:00 UTC arXiv cs.AI S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF