LLMs for medical consultation evaluated after problem framing; study analyzes three API models across multi-turn vignettes and standardized-patient simulations
Read the original at arxiv.org→arXiv:2608.17330v1 Announce Type: new Abstract: Large language models for medical consultation are often evaluated after a clinical problem has already been made clear, although real consultations may begin with a...
Original headline: "LLMs for Medical Consultation Are Evaluated Too Late: The Preformulation Gap"
Coverage timeline
- Aug 19, 04:00 UTC arXiv cs.AI lead source LLMs for Medical Consultation Are Evaluated Too Late: The Preformulation Gap