Diagnosing correctness probes under self-judgement confounding
Read the original at arxiv.org→arXiv:2607.16799v1 Announce Type: new Abstract: Hidden-state readouts can predict whether language-model outputs are correct, but objective correctness (OC) usually agrees with the model's own self-judgement (SJ),...
Original headline: "Diagnosing Correctness Probes under Self-Judgement Confounding"