Harmfulness propagation dynamics: layer-wise trajectories of adversarial intent in large language models
Read the original at arxiv.org→arXiv:2609.13534v1 Announce Type: new Abstract: We identify \textbf{Harmfulness Propagation Dynamics (HPD)}: for harmful prompts, the projection of the last-token hidden state onto a learned harm direction rises...
Original headline: "Harmfulness Propagation Dynamics: Layer-wise Trajectories of Adversarial Intent in Large Language Models"
Coverage timeline
- Sep 15, 04:00 UTC arXiv cs.CL lead source Harmfulness Propagation Dynamics: Layer-wise Trajectories of Adversarial Intent in Large Language Models