Evaluating LLM reliability beyond accuracy: how model answers vary with meaning-preserving paraphrases across tasks
Read the original at arxiv.org→arXiv:2607.22554v1 Announce Type: new Abstract: Large language models (LLMs) often achieve strong accuracy on benchmarks, yet it remains unclear how reliably they apply this knowledge when the same question is...
Original headline: "Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy"
Coverage timeline
- Jul 28, 04:00 UTC arXiv cs.AI lead source Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy