When do internal probes beat reading the answer? Miscalibrated readouts and behavior-concealed knowledge in language models
Read the original at arxiv.org→arXiv:2609.04582v1 Announce Type: new Abstract: A 0.6B language model, asked to verify 1,200 logical conclusions (half valid, half corrupted by a single semantic edit), answers YES every time. Judged by behavior it...
Original headline: "When Do Internal Probes Beat Reading the Answer? Miscalibrated Readouts and Behavior-Concealed Knowledge in Language Models"
Coverage timeline
- Sep 7, 04:00 UTC arXiv cs.CL lead source When Do Internal Probes Beat Reading the Answer? Miscalibrated Readouts and Behavior-Concealed Knowledge in Language Models