From detection to refusal: safer LLMs via circuit-guided weight scaling
Read the original at arxiv.org→arXiv:2609.00051v1 Announce Type: new Abstract: Despite extensive alignment efforts, Large Language Models (LLMs) remain vulnerable to generating unsafe content under adversarial prompting, yet the internal...
Original headline: "From Detection to Refusal: Safer LLMs via Circuit-Guided Weight Scaling"
Coverage timeline
- Sep 2, 04:00 UTC arXiv cs.CL lead source From Detection to Refusal: Safer LLMs via Circuit-Guided Weight Scaling