LLMs temper Bayesian priors using a single unembedding direction, the direction of ignorance
Read the original at arxiv.org→arXiv:2609.02959v1 Announce Type: new Abstract: What does a language model predict when it has few clues? The answer lurks in its unembedding geometry: a single direction of the unembedding matrix encodes the...
Original headline: "The Geometry of Ignorance: LLMs Know When to Temper Bayesian Priors"
Coverage timeline
- Sep 4, 04:00 UTC arXiv cs.LG lead source The Geometry of Ignorance: LLMs Know When to Temper Bayesian Priors