KlinikeBench evaluates language models beyond diagnostic accuracy by assessing how well models gather information and determine appropriate clinical assessments.
Read the original at arxiv.org→arXiv:2609.38480v1 Announce Type: new Abstract: Most clinical benchmarks evaluate language models (LMs) on diagnosis using complete case descriptions. In clinical practice, however, patients present information in...
Original headline: "KlinikeBench: Evaluating Language Models Beyond Diagnostic Accuracy"
Coverage timeline
- Oct 1, 04:00 UTC arXiv cs.CL lead source KlinikeBench: Evaluating Language Models Beyond Diagnostic Accuracy