Safety, or just capability? A validity audit of agent-safety benchmarks
Read the original at arxiv.org→arXiv:2607.28685v1 Announce Type: new Abstract: Agent-safety benchmarks measure different behaviors, and their scores get quoted interchangeably as an agent's safety. We treat four of them (R-Judge, InjecAgent,...
Original headline: "Safety, or Just Capability? A Validity Audit of Agent-Safety Benchmarks"
Coverage timeline
- Jul 31, 06:46 UTC Hacker News (AI) Benchmarking Guardrails for AI Agent Safety
- Aug 3, 04:00 UTC arXiv cs.AI lead source Safety, or Just Capability? A Validity Audit of Agent-Safety Benchmarks