Component and dimension sparsity in transformer refusal mechanisms
Read the original at arxiv.org→arXiv:2610.06903v1 Announce Type: new Abstract: Activation steering manipulates large language model behavior by intervening on internal activations, but the mechanistic basis of these interventions remains poorly...
Original headline: "Component and Dimension Sparsity in Transformer Refusal Mechanisms"
Coverage timeline
- Oct 7, 04:00 UTC arXiv cs.CL lead source Component and Dimension Sparsity in Transformer Refusal Mechanisms