When Noise Fabricates Bias: the fragility of LLM-as-a-judge bias measurement under noisy text
Read the original at arxiv.org→arXiv:2609.11067v1 Announce Type: new Abstract: Large language models are increasingly used as judges to measure social bias in text, yet the passages they judge are often noisy, containing typos, informal spelling,...
Original headline: "When Noise Fabricates Bias: The Fragility of LLM-as-a-Judge Bias Measurement under Noisy Text"
Coverage timeline
- Sep 11, 04:00 UTC arXiv cs.CL lead source When Noise Fabricates Bias: The Fragility of LLM-as-a-Judge Bias Measurement under Noisy Text