GAUGE shows that offline LLM judges may misrank task-oriented agents; evaluation uses persona-driven user simulators and a grounded verifiable reward across 25 agents from six providers
Read the original at arxiv.org→arXiv:2609.12191v1 Announce Type: new Abstract: Comparing and selecting task-oriented LLM agents increasingly relies on a low-cost offline evaluation gate: persona-driven LLM user-simulators converse with each...
Original headline: "GAUGE: When Not to Trust LLM-as-a-Judge in User-Simulated Evaluation of Task-Oriented Agents"
Coverage timeline
- Sep 14, 04:00 UTC arXiv cs.CL lead source GAUGE: When Not to Trust LLM-as-a-Judge in User-Simulated Evaluation of Task-Oriented Agents
- Sep 14, 04:00 UTC arXiv cs.LG Can We Trust LLM Judges: A Study of Capability-Dependent Biases and Multi-Judge Ensemble for Bias Calibration