ArcticQA: a dataset of 194 Arctic science questions with automated checks of answer support; ArcticAbstain benchmark compares answer presence versus abstention by LLMs
Read the original at arxiv.org→arXiv:2610.09446v1 Announce Type: new Abstract: Large language models (LLMs) should abstain from scientific multiple-choice questions when no option is valid, but frequent abstention alone does not demonstrate...
Original headline: "Arctic Questions, Missing Answers: A Dataset and Benchmark for LLM Abstention in Arctic Science"
Coverage timeline
- Oct 8, 04:00 UTC arXiv cs.CL lead source Arctic Questions, Missing Answers: A Dataset and Benchmark for LLM Abstention in Arctic Science