Graded entity-familiarity readouts in language models: Polish adaptation, cross-language robustness, and refusal steering
Read the original at arxiv.org→arXiv:2607.13568v1 Announce Type: new Abstract: Can a language model estimate its familiarity with an entity before generating an answer? We study activations at the final prompt token in twelve instruction-tuned...
Original headline: "Graded Entity-Familiarity Readouts in Language Models: Polish Adaptation, Cross-Language Robustness, and Refusal Steering"
Coverage timeline
- Jul 16, 04:00 UTC arXiv cs.CL lead source Graded Entity-Familiarity Readouts in Language Models: Polish Adaptation, Cross-Language Robustness, and Refusal Steering