Caution for the Environment: Multimodal LLM Agents are Susceptible to Environmental Distractions
Xinbei Ma, Yiting Wang, Yao Yao, Tongxin Yuan, Aston Zhang, Zhuosheng Zhang, Hai Zhao
Abstract
This paper investigates the faithfulness of multimodal large language model (MLLM) agents in a graphical user interface (GUI) environment, aiming to address the research question of whether multimodal GUI agents can be distracted by environmental context. A general scenario is proposed where both the user and the agent are benign, and the environment, while not malicious, contains unrelated contents. A wide range of MLLMs are evaluated as GUI agents using a simulated dataset, following three working patterns with different levels of perception. Experimental results reveal that even the most powerful models, whether generalist agents or specialist GUI agents, are susceptible to distractions. While recent studies predominantly focus on the helpfulness of agents, our findings first indicate that these agents are prone to environmental distractions. Furthermore, we implement an adversarial environment injection and analyze the approach to improve faithfulness, calling for a collective focus on this important topic. The code is available at https://github.com/xbmxb/EnvDistraction .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ddf6a3d3-8de2-43c1-a3ed-c24bf6d323f6Cited by top-tier papers10
- macOSWorld: A Multilingual Interactive Benchmark for GUI AgentsPei Yang, Hai Ci, Mike Zheng ShouNeurIPS 2025 · 34 citations
- RiOSWorld: Benchmarking the Risk of Multimodal Computer-Use AgentsJingyi Yang, Shuai Shao, Dongrui Liu, Jing ShaoNeurIPS 2025 · 33 citations
- UltraHorizon: Benchmarking LLM-Agent Capabilities in Ultra Long-Horizon ScenariosHaotian Luo, Huaisong Zhang, Xuelin Zhang, Haoyu Wang et al.ICML 2026 · 21 citations
- OS-Sentinel: Towards Safety-Enhanced Mobile GUI Agents via Hybrid Validation in Realistic WorkflowsQiushi Sun, Mukai Li, Zhoumianze Liu, Zhihui Xie et al.ACL 2026 · 14 citations
- Are Large Language Models Sensitive to the Motives Behind Communication?Addison J. Wu, Ryan Liu, Kerem Oktar, Theodore R. Sumers et al.NeurIPS 2025 · 9 citations
Builds on22
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
- Large Language Models Can Be Easily Distracted by Irrelevant ContextFreda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales et al.ICML 2023 · 970 citations
Related papers
- LaSM: Layer-wise Scaling Mechanism for Defending Pop-up Attack on GUI AgentsZihe Yan, Zhuosheng Zhang, Jiaping Gui, Gongshen LiuCVPR 2026 · 5 citations
- GUI-World: A Video Benchmark and Dataset for Multimodal GUI-oriented UnderstandingDongping Chen, Yue Huang, Siyuan Wu, Jingyu Tang et al.ICLR 2025 · 1 citation
- MP-GUI: Modality Perception with MLLMs for GUI UnderstandingZiwei Wang, Weizhi Chen, Leyang Yang, Sheng Zhou et al.CVPR 2025
- PG-Agent: An Agent Powered by Page GraphWeizhi Chen, Ziwei Wang, Leyang Yang, Sheng Zhou et al.ACM MM 2025
- Environmental Injection Attacks against GUI Agents in Realistic Dynamic EnvironmentsYitong Zhang, Ximo Li, Liyi Cai, Jia LiISSTA 2026
