K-Bench: evaluating model performance on real, underspecified scientific agent requests from live user traffic
Read the original at arxiv.org→arXiv:2608.21601v1 Announce Type: new Abstract: Benchmarks for scientific artificial intelligence are mostly written to be scored: multiple-choice questions, curated agent tasks with reference solutions, or...
Original headline: "K-Bench: measuring model performance on real scientific agent requests"
Coverage timeline
- Aug 25, 04:00 UTC arXiv cs.AI lead source K-Bench: measuring model performance on real scientific agent requests