From token probabilities to calibrated confidence: an empirical study of mathematical question answering
Read the original at arxiv.org→arXiv:2608.07827v1 Announce Type: new Abstract: Confidence estimation for large language models (LLMs) aims to estimate the probability that a generated answer is correct, while calibration aligns these estimates...
Original headline: "From token probabilities to calibrated confidence: An empirical study of mathematical question answering"
Coverage timeline
- Aug 11, 04:00 UTC arXiv cs.LG lead source From token probabilities to calibrated confidence: An empirical study of mathematical question answering