Typos can rotate LLM readout vectors and undermine prompt-detecting probes, with small edits persisting for several tokens; three common typos per message can impact a single-position prompt-injection, according to arXiv:2609.15994v1.
Read the original at arxiv.org→arXiv:2609.15994v1 Announce Type: new Abstract: LLMs handle ordinary typing variation fluently: a typo or missing punctuation leaves both user intent and the model's response substantively unchanged. Yet probes that...
Original headline: "Latent Undertow: How Ordinary Typos Break Probes"
Coverage timeline
- Sep 16, 04:00 UTC arXiv cs.CL lead source Latent Undertow: How Ordinary Typos Break Probes