SciLitBench: benchmark and design principles for LLM-powered systematic literature reviews
Read the original at arxiv.org→arXiv:2609.05505v1 Announce Type: new Abstract: Systematic reviews require sustained human judgment across thousands of records, yet existing evaluations of large language models (LLMs) typically examine review...
Original headline: "SciLitBench: Benchmark and Design Principles for LLM-Powered Systematic Literature Reviews"
Coverage timeline
- Sep 9, 04:00 UTC arXiv cs.AI lead source SciLitBench: Benchmark and Design Principles for LLM-Powered Systematic Literature Reviews