Evaluating language model safety across long adversarial conversations
Read the original at arxiv.org→arXiv:2609.38357v1 Announce Type: new Abstract: Conversational safety evaluations often test language models with a single harmful prompt, even though real-world systems interact with users through long, adaptive...
Original headline: "Evaluating Language Model Safety Across Long Adversarial Conversations"
Coverage timeline
- Oct 1, 04:00 UTC arXiv cs.CL lead source Evaluating Language Model Safety Across Long Adversarial Conversations