There is no neutral harness: modern LLM leaderboards are affected by harness sensitivity in item variance
Read the original at arxiv.org→arXiv:2608.21382v1 Announce Type: new Abstract: Multiple-choice benchmarks fix the questions and the correct answers, but not the harness: the order of the options, the wording of the prompt, and whether a language...
Original headline: "There Is No Neutral Harness: Modern LLM Leaderboards Are Manufactured by Config-Fragile Items"
Coverage timeline
- Aug 25, 04:00 UTC arXiv cs.AI lead source There Is No Neutral Harness: Modern LLM Leaderboards Are Manufactured by Config-Fragile Items