Refusal is not robustness: auditing confident fabrication in large language models on a provably uninformative clinical pain speech transcript
Read the original at arxiv.org→arXiv:2608.26167v1 Announce Type: new Abstract: Hallucination and abstention benchmarks rarely establish that a model could not have known the correct answer, making it difficult to distinguish appropriate...
Original headline: "Refusal Is Not Robustness: Auditing Confident Fabrication in Large Language Models on a Provably Uninformative Clinical Pain Speech Transcript"
Coverage timeline
- Aug 28, 04:00 UTC arXiv cs.AI lead source Refusal Is Not Robustness: Auditing Confident Fabrication in Large Language Models on a Provably Uninformative Clinical Pain Speech Transcript