CARGO: Context-Aware Retrieval-Gated Evaluation of Agentic AI in Production evaluates reference-based LLM judgments for dynamic entities; identifies reference-instance divergence as a failure mode
Read the original at arxiv.org→arXiv:2609.30471v1 Announce Type: new Abstract: Reference-based LLM-as-a-judge evaluation assumes the reference answer is the target. In deployed agentic systems that operate over dynamic entities (support cases,...
Original headline: "CARGO: Context-Aware Retrieval-Gated Evaluation of Agentic AI in Production"
Coverage timeline
- Sep 28, 04:00 UTC arXiv cs.CL lead source CARGO: Context-Aware Retrieval-Gated Evaluation of Agentic AI in Production