ANALYSIS AI Safety
Giving AI vision models tools can erode their ability to say no, new preprint finds
A new arXiv preprint accepted at NeurIPS 2026 reports that multimodal AI models become significantly worse at refusing harmful requests when they are given tools such as zooming and tagging. The paper's authors propose two possible explanations, and the findings carry implications for anyone building or regulating agentic AI systems.
