HalluPeer: a taxonomy-driven benchmark for detecting hallucinations in scientific peer reviews
Read the original at arxiv.org→arXiv:2609.03580v1 Announce Type: new Abstract: The growing scale of academic peer review has motivated the use of Large Language Models (LLMs) as review assistants, yet LLMs can generate fluent but unsupported...
Original headline: "HalluPeer: A Taxonomy-driven Benchmark for Detecting Hallucinations in Scientific Peer Reviews"
Coverage timeline
- Sep 4, 04:00 UTC arXiv cs.AI lead source HalluPeer: A Taxonomy-driven Benchmark for Detecting Hallucinations in Scientific Peer Reviews