Safe-role internalization improves robustness and generalization of LLM safety alignment
Read the original at arxiv.org→arXiv:2610.07023v1 Announce Type: new Abstract: Large Language Models (LLMs) have achieved remarkable capabilities but remain vulnerable to jailbreak attacks that elicit harmful or unsafe outputs. Existing safety...
Original headline: "Beyond Refusal Patterns: Safe-Role Internalization for Robust and Generalizable LLM Safety Alignment"
Coverage timeline
- Oct 7, 04:00 UTC arXiv cs.AI lead source Beyond Refusal Patterns: Safe-Role Internalization for Robust and Generalizable LLM Safety Alignment