MAWILE: Multi-Axis Workbench for Inspecting LLM Evaluators
Read the original at arxiv.org→arXiv:2609.22599v1 Announce Type: new Abstract: Large language model (LLM) judges provide a flexible and scalable method for evaluating model and agent outputs, but their verdicts can be sensitive to incidental...
Coverage timeline
- Sep 22, 04:00 UTC arXiv cs.AI lead source MAWILE: Multi-Axis Workbench for Inspecting LLM Evaluators