Do large language models scrutinise what they review? A multimodal audit of scoring calibration, error detection, and author-identity effects
Read the original at arxiv.org→arXiv:2608.28626v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to generate peer reviews, prompting examination of their capacity for critical evaluation. This study evaluates two...
Coverage timeline
- Sep 1, 04:00 UTC arXiv cs.CL lead source Do large language models scrutinise what they review? A multimodal audit of scoring calibration, error detection, and author-identity effects