Are we grading properly? Understanding failure modes in medical benchmarks
Read the original at arxiv.org→arXiv:2609.16023v1 Announce Type: new Abstract: Medical evaluation is shifting from static option-based questioning to realistic clinical scenarios with open-ended output modes. Grading these at scale naively,...
Original headline: "Are We Grading Properly? Understanding Failure Modes in Medical Benchmarks"
Coverage timeline
- Sep 16, 04:00 UTC arXiv cs.CL lead source Are We Grading Properly? Understanding Failure Modes in Medical Benchmarks