Auditing LLM judges for occupational AI measurement reveals scale issues; O*NET-BENCH audit suite evaluated across model families shows 25 configurations achieve tie-aware agreement across ratings
Read the original at arxiv.org→arXiv:2610.02492v1 Announce Type: new Abstract: LLM judges are increasingly used to assess whether AI outputs meet workplace requirements, but agreement on response rankings does not establish agreement on...
Original headline: "Right Order, Wrong Scale: Auditing LLM Judges for Occupational AI Measurement"
Coverage timeline
- Oct 5, 04:00 UTC arXiv cs.AI lead source Right Order, Wrong Scale: Auditing LLM Judges for Occupational AI Measurement