Sparse readout prism: explaining logit-lens scores in features instead of tokens
Read the original at arxiv.org→arXiv:2609.01936v1 Announce Type: new Abstract: A language model's prediction of its next token develops across layers, and lens methods track this process by decoding intermediate hidden states into tokens. But a...
Original headline: "Sparse Readout Prism: Explaining Logit-Lens Scores in Features Instead of Tokens"
Coverage timeline
- Sep 3, 04:00 UTC arXiv cs.CL lead source Sparse Readout Prism: Explaining Logit-Lens Scores in Features Instead of Tokens