Clinician use of language models differs from how the models are evaluated
Read the original at arxiv.org→arXiv:2610.11069v1 Announce Type: new Abstract: Large language model (LLM) assistants are being deployed to clinicians across health systems, and judgments about their readiness rest largely on benchmark scores,...
Original headline: "Clinician use of language models diverges from how the models are evaluated"
Coverage timeline
- Oct 9, 04:00 UTC arXiv cs.CL lead source Clinician use of language models diverges from how the models are evaluated