AgentNoiseBench: Benchmarking Robustness of Tool-Using LLM Agents Under Noisy Condition
Ruipeng Wang, Yuxin Chen, Yukai Wang, Chang Wu, Junfeng Fang, Xiaodong Cai, Qi GU, Hui Su, An Zhang, Xiang Wang, Xunliang Cai, Tat-Seng Chua
摘要
Recent advances in large language models have enabled LLM-based agents to achieve strong performance on a variety of benchmarks. However, their performance in real-world deployments often that observed on benchmark settings, especially in complex and imperfect environments. This discrepancy largely arises because prevailing training and evaluation paradigms are typically built on idealized assumptions, overlooking the inherent stochasticity and noise present in real-world interactions. To bridge this gap, we introduce AgentNoiseBench, a framework for systematically evaluating the robustness of agentic models under noisy environments. We first conduct an in-depth analysis of biases and uncertainties in real-world scenarios and categorize environmental noise into two primary types: usernoise and tool-noise. Building on this analysis, we develop an automated pipeline that injects controllable noise into existing agent-centric benchmarks while preserving task solvability. Leveraging this pipeline, we perform extensive evaluations across a wide range of models with diverse architectures and parameter scales. Our results reveal consistent performance variations under different noise conditions, highlighting the sensitivity of current agentic models to realistic environmental perturbations. Code is available at https://github.com/keven-cyber/ agentnoisebench
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu 等NeurIPS 2023 · 被引用 5,989 次
- MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to UseYue Huang, Jiawen Shi, Yuan Li, Chenrui Fan 等ICLR 2024 · 被引用 188 次
- Evaluating the Instruction-Following Robustness of Large Language Models to Prompt InjectionZekun Li, Baolin Peng, Pengcheng He, Xifeng YanEMNLP 2024 · 被引用 15 次
- Scenario-independent Uncertainty Estimation for LLM-based Question Answering via Factor AnalysisZhihua Wen, Zhizhao Liu, Zhiliang Tian, Shilong Pan 等WWW 2025 · 被引用 7 次
- ToolBeHonest: A Multi-level Hallucination Diagnostic Benchmark for Tool-Augmented Large Language ModelsYuxiang Zhang, Jing Chen, Junjie Wang, Yaxin Liu 等EMNLP 2024 · 被引用 6 次
相关 Paper
- AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World ContextsKeyu Li, Junhao Shi, Yang Xiao, Mohan Jiang 等ACL 2026 · 被引用 14 次
- AgentGym2: Benchmarking Large Language Model Agents in De-Idealized Real-World EnvironmentsZhiheng Xi, Dingwen Yang, Jiaqi Liu, Jixuan Huang 等ACL 2026
- Pandora's Box or Aladdin's Lamp: A Comprehensive Analysis Revealing the Role of RAG Noise in Large Language ModelsJinyang Wu, Shuai Zhang, Feihu Che, Mingkuan Feng 等ACL 2025 · 被引用 12 次
- Learning to Ask: When LLM Agents Meet Unclear InstructionWenxuan Wang, Juluan Shi, Zixuan Ling, Yuk-Kit Chan 等EMNLP 2025 · 被引用 1 次
- RoTBench: A Multi-Level Benchmark for Evaluating the Robustness of Large Language Models in Tool LearningJunjie Ye, Yilong Wu, Songyang Gao, Caishuang Huang 等EMNLP 2024 · 被引用 6 次
