RePro: proof-verified benchmark rewriting for reliable evaluation of LLM mathematical problem solving
Read the original at arxiv.org→arXiv:2609.00062v1 Announce Type: new Abstract: Data contamination undermines the reliable evaluation of large language models (LLMs) on mathematical problem solving. While rewriting-based evaluation mitigates...
Original headline: "RePro: Proof-Verified Benchmark Rewriting for Reliable Evaluation of LLM Mathematical Problem Solving"
Coverage timeline
- Sep 2, 04:00 UTC arXiv cs.CL lead source RePro: Proof-Verified Benchmark Rewriting for Reliable Evaluation of LLM Mathematical Problem Solving