Codifying the Judge: scalable evaluation via program distillation
Read the original at arxiv.org→arXiv:2607.22561v1 Announce Type: new Abstract: LLM-as-a-judge has become the standard for automated evaluation, but it suffers from high cost, significant latency, and opaque decisions -- limitations that undermine...
Original headline: "Codifying the Judge: Scalable Evaluation via Program Distillation"
Coverage timeline
- Jul 28, 04:00 UTC arXiv cs.AI lead source Codifying the Judge: Scalable Evaluation via Program Distillation