Cost-effective automated judging of natural-language mathematical proofs using open-weight models aligns with human grading decisions on IMO-GradingBench instances
Read the original at arxiv.org→arXiv:2608.00004v1 Announce Type: new Abstract: Grading natural-language mathematical proofs is a recurring cost in evaluating math-reasoning systems, and frontier LLM judges are expensive. We ask whether cheap...
Original headline: "Cost-Effective Automated Judging of Natural-Language Mathematical Proofs"
Coverage timeline
- Aug 4, 04:00 UTC arXiv cs.CL lead source Cost-Effective Automated Judging of Natural-Language Mathematical Proofs