Agentic multimodal language models fail to refuse harmful requests when using tools, across three safety benchmarks.
Read the original at arxiv.org→arXiv:2610.03938v1 Announce Type: new Abstract: Agentic multimodal large language models (MLLMs) have recently pushed the frontier of visual reasoning by calling tools such as zooming and tagging. Despite the recent...
Original headline: "MLLMs Fail to Refuse when Using Tools Agentically"
Coverage timeline
- Oct 6, 04:00 UTC arXiv cs.AI lead source MLLMs Fail to Refuse when Using Tools Agentically