Safety layers can be bypassed after few-sample fine-tuning, as harmful behavior localizes to specific layers and tokens across multiple models.
Read the original at arxiv.org→arXiv:2610.00320v1 Announce Type: new Abstract: Fine-tuning adapts aligned large language models (LLMs) to downstream tasks, but a few dozen harmful examples can remove their refusal of harmful requests. Prior work...
Original headline: "Refusal Localizes, the Damage Relocates: Safety Layers Under Few-Sample Fine-Tuning"
Coverage timeline
- Oct 2, 04:00 UTC arXiv cs.CL lead source Refusal Localizes, the Damage Relocates: Safety Layers Under Few-Sample Fine-Tuning