Auditing and repairing LLM-as-judge failures in a production text-to-SQL pipeline
Read the original at arxiv.org→arXiv:2609.30290v1 Announce Type: new Abstract: Production text-to-SQL pipelines often end with an LLM-as-judge whose agreement with human annotators has never actually been measured. When we checked ours, the...
Original headline: "Auditing and Repairing LLM-as-Judge Failures in a Production Text-to-SQL Pipeline"
Coverage timeline
- Sep 28, 04:00 UTC arXiv cs.CL lead source Auditing and Repairing LLM-as-Judge Failures in a Production Text-to-SQL Pipeline