Recall is not protection: evaluating safety monitors against model compliance
Read the original at arxiv.org→arXiv:2609.05797v1 Announce Type: new Abstract: Safety monitors screen prompts sent to deployed language models, flagging harmful requests so they are never answered. They are evaluated by recall against harmfulness...
Original headline: "Recall Is Not Protection: Evaluating Safety Monitors Against Model Compliance"
Coverage timeline
- Sep 9, 04:00 UTC arXiv cs.CL lead source Recall Is Not Protection: Evaluating Safety Monitors Against Model Compliance