No usable linear "capitulation direction" in two small LLMs: a validation protocol for activation-steering claims and a cross-family study of sycophancy under pushback
Read the original at arxiv.org→arXiv:2609.17550v1 Announce Type: new Abstract: Language models frequently abandon correct answers when users push back. We study this in two small instruction-tuned models from different families, Qwen2.5-1.5B and...
Original headline: "No Usable Linear "Capitulation Direction" in Two Small LLMs: A Validation Protocol for Activation-Steering Claims, and a Cross-Family Behavioral Study of Sycophancy Under Pushback"