When Is a multi-agent code judge actually grounded? Two label-free measurements, and a judge that declines to guess
Read the original at arxiv.org→arXiv:2609.30328v1 Announce Type: new Abstract: When one language model judges whether another's code is correct, it does not report the absence of evidence. It returns a confident verdict with reasoning attached,...
Original headline: "When Is a Multi-Agent Code Judge Actually Grounded? Two Label-Free Measurements, and a Judge That Declines to Guess"
Coverage timeline
- Sep 28, 04:00 UTC arXiv cs.AI lead source When Is a Multi-Agent Code Judge Actually Grounded? Two Label-Free Measurements, and a Judge That Declines to Guess