AgentNoiseBench: Benchmarking Robustness of Tool-Using LLM Agents Under Noisy Condition
Ruipeng Wang, Yuxin Chen, Yukai Wang, Chang Wu, Junfeng Fang, Xiaodong Cai, Qi GU, Hui Su, An Zhang, Xiang Wang, Xunliang Cai, Tat-Seng Chua
Abstract
Recent advances in large language models have enabled LLM-based agents to achieve strong performance on a variety of benchmarks. However, their performance in real-world deployments often that observed on benchmark settings, especially in complex and imperfect environments. This discrepancy largely arises because prevailing training and evaluation paradigms are typically built on idealized assumptions, overlooking the inherent stochasticity and noise present in real-world interactions. To bridge this gap, we introduce AgentNoiseBench, a framework for systematically evaluating the robustness of agentic models under noisy environments. We first conduct an in-depth analysis of biases and uncertainties in real-world scenarios and categorize environmental noise into two primary types: usernoise and tool-noise. Building on this analysis, we develop an automated pipeline that injects controllable noise into existing agent-centric benchmarks while preserving task solvability. Leveraging this pipeline, we perform extensive evaluations across a wide range of models with diverse architectures and parameter scales. Our results reveal consistent performance variations under different noise conditions, highlighting the sensitivity of current agentic models to realistic environmental perturbations. Code is available at https://github.com/keven-cyber/ agentnoisebench
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 069475ae-80e7-44fd-a858-e68c460b00eeBuilds on13
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu et al.NeurIPS 2023 · 5,989 citations
- MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to UseYue Huang, Jiawen Shi, Yuan Li, Chenrui Fan et al.ICLR 2024 · 188 citations
- Evaluating the Instruction-Following Robustness of Large Language Models to Prompt InjectionZekun Li, Baolin Peng, Pengcheng He, Xifeng YanEMNLP 2024 · 15 citations
- Scenario-independent Uncertainty Estimation for LLM-based Question Answering via Factor AnalysisZhihua Wen, Zhizhao Liu, Zhiliang Tian, Shilong Pan et al.WWW 2025 · 7 citations
- ToolBeHonest: A Multi-level Hallucination Diagnostic Benchmark for Tool-Augmented Large Language ModelsYuxiang Zhang, Jing Chen, Junjie Wang, Yaxin Liu et al.EMNLP 2024 · 6 citations
Related papers
- AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World ContextsKeyu Li, Junhao Shi, Yang Xiao, Mohan Jiang et al.ACL 2026 · 14 citations
- AgentGym2: Benchmarking Large Language Model Agents in De-Idealized Real-World EnvironmentsZhiheng Xi, Dingwen Yang, Jiaqi Liu, Jixuan Huang et al.ACL 2026
- Pandora's Box or Aladdin's Lamp: A Comprehensive Analysis Revealing the Role of RAG Noise in Large Language ModelsJinyang Wu, Shuai Zhang, Feihu Che, Mingkuan Feng et al.ACL 2025 · 12 citations
- Learning to Ask: When LLM Agents Meet Unclear InstructionWenxuan Wang, Juluan Shi, Zixuan Ling, Yuk-Kit Chan et al.EMNLP 2025 · 1 citation
- RoTBench: A Multi-Level Benchmark for Evaluating the Robustness of Large Language Models in Tool LearningJunjie Ye, Yilong Wu, Songyang Gao, Caishuang Huang et al.EMNLP 2024 · 6 citations
