How much of a measured AI preference is the model and how much is the instrument; multiple instruments disagree on the source of model preferences
Read the original at arxiv.org→arXiv:2608.23641v1 Announce Type: new Abstract: Model welfare research infers what a model prefers from the answers returned to prompts written to elicit preferences. Keeling et al. (2024), Mazeika et al. (2025),...
Original headline: "How much of a measured AI preference is the model, and how much is the instrument?"
Coverage timeline
- Aug 26, 04:00 UTC arXiv cs.AI lead source How much of a measured AI preference is the model, and how much is the instrument?