SHELF: a synthetic harness for evaluating LLM fitness in multi-task bibliographic benchmarking
Read the original at arxiv.org→arXiv:2609.03047v1 Announce Type: new Abstract: Libraries and archives manage large collections with limited staff and computing budgets, yet common benchmarks do not systematically test their bibliographic work....
Original headline: "SHELF: A Synthetic Harness for Multi-Task Bibliographic Benchmarking"
Coverage timeline
- Sep 4, 04:00 UTC arXiv cs.CL lead source SHELF: A Synthetic Harness for Multi-Task Bibliographic Benchmarking