Same patient, different order: action-level reliability of clinical LLM agents under repeated runs
Read the original at arxiv.org→arXiv:2609.13582v1 Announce Type: new Abstract: A clinical agent benchmark can report the same verdict on identical inputs while the agent files a materially different order on each run. Such agents order tests,...
Original headline: "Same Patient, Different Order: Action-Level Reliability of Clinical LLM Agents Under Repeated Runs"
Coverage timeline
- Sep 15, 04:00 UTC arXiv cs.CL lead source Same Patient, Different Order: Action-Level Reliability of Clinical LLM Agents Under Repeated Runs