No One Model Catches Every Harm: Benchmarking Content Moderation Across Safety Scenarios
Read the original at arxiv.org→arXiv:2608.21775v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed in real-world applications, yet they remain vulnerable to generating harmful content. From adversarial...
Coverage timeline
- Aug 25, 04:00 UTC arXiv cs.CL lead source No One Model Catches Every Harm: Benchmarking Content Moderation Across Safety Scenarios