SAFEGuard detects optimization-based jailbreak attacks through harmful semantic analysis and fluency measurement
Read the original at arxiv.org→arXiv:2609.05850v1 Announce Type: new Abstract: Despite the significant efforts devoted to aligning large language models (LLMs) with human values and ensuring safe deployment, recent work has revealed that LLMs...
Original headline: "SAFEGuard: Detect Optimization-Based Jailbreak Attacks Through Harmful Semantic Analysis and Fluency Measurement"
Coverage timeline
- Sep 9, 04:00 UTC arXiv cs.LG lead source SAFEGuard: Detect Optimization-Based Jailbreak Attacks Through Harmful Semantic Analysis and Fluency Measurement