LLM-as-a-Judge: rubric text alone predicts judge outputs in automated text-generation evaluation
Read the original at arxiv.org→arXiv:2609.02942v1 Announce Type: new Abstract: LLM-as-a-Judge pipelines are increasingly used to evaluate AI-generated text, based on the assumption that judgments arise from reasoning over candidate responses with...
Original headline: "Judging LLM-as-a-Judge: Concerning Rubric Artifacts in LLM-based Automated Text Generation Evaluation"
Coverage timeline
- Sep 4, 04:00 UTC arXiv cs.CL lead source Judging LLM-as-a-Judge: Concerning Rubric Artifacts in LLM-based Automated Text Generation Evaluation