Didactic data vs. clinical data shape medical large language models; token-matched experiments show how data composition affects performance and error patterns
Read the original at arxiv.org→arXiv:2609.22161v1 Announce Type: new Abstract: Medical large language models are commonly trained on mixtures of didactic data (e.g., textbooks) and clinical data (e.g., patient records), yet how these data types...
Original headline: "Didactic knowledge or Clinical Cases? How Data Types Shape Medical Large Language Models"
Coverage timeline
- Sep 22, 04:00 UTC arXiv cs.AI lead source Didactic knowledge or Clinical Cases? How Data Types Shape Medical Large Language Models