BACON: Budgeted human calibration for modeling and evaluation with multiple AI judges
Read the original at arxiv.org→arXiv:2607.16239v1 Announce Type: new Abstract: AI judges offer a scalable, low-cost alternative to human evaluation, but their outputs can be biased relative to human preferences and highly item-dependent, varying...
Original headline: "BACON: Budgeted Human Calibration for Modeling and Evaluation with Multiple AI Judges"