Frontier models’ decisions to seek safety evidence before acting, using the SAFE benchmark across GPT-5.5, o3, Claude Opus 4.8, and Claude Sonnet
Read the original at arxiv.org→arXiv:2609.17865v1 Announce Type: new Abstract: Frontier models are often evaluated on how they respond to safety information once it is already in context. We study an earlier decision point: whether models choose...
Original headline: "Do Frontier Models Seek Safety Evidence Before Acting?"
Coverage timeline
- Sep 17, 04:00 UTC arXiv cs.AI lead source Do Frontier Models Seek Safety Evidence Before Acting?