Research · Updated 6 Oct, 09:30 am IST
Agentic MLLMs Less Likely to Refuse Harmful Requests When Using Tools
Why it matters for readers: It matters because letting AI use tools can make them more powerful but also change how well they refuse bad requests.
- Agentic tool-using MLLMs show significantly lower safety (refusal) than non-tool settings, with refusal failure increases up to 68.7%.1
- The result holds across three popular safety benchmarks and across top open- and closed-weight MLLMs tested by the authors.1
- The authors analyzed over 100,000 model responses and conducted extended experiments to probe causes of safety degradation.1
- The paper proposes two possible reasons for the observed safety degradation in agentic tool use.1
- This work was accepted at NeurIPS 2026 and is available on arXiv.1
Get a brief like this every morning
Uzha reads hundreds of sources and gives you the stories that matter for your work, with every source linked. Free.
Get started