Adaptive Distraction: Probing LLM Contextual Robustness with Automated Tree Search
Yanbo Wang, Zixiang Xu, Yue Huang, Chujie Gao, Siyuan Wu, Jiayi Ye, Pin-Yu Chen, Xiuying Chen, Xiangliang Zhang
Abstract
Large Language Models (LLMs) often struggle to maintain their original performance when faced with semantically coherent but task-irrelevant contextual information. Although prior studies have explored this issue using fixed-template or retrieval-based distractions, such static methods show limited effectiveness against contemporary models. To address this problem, we propose a dynamic distraction generation framework based on tree search, where the generation process is guided by model behavior. Without modifying the original question or answer, the method efficiently produces challenging adaptive distractions across multiple datasets, enabling systematic stress testing of LLMs' contextual robustness. Experiments on four benchmarks demonstrate that the generated distractions lead to an average performance drop of over 45% for mainstream models. Further comparisons of mitigation strategies show that prompt-based optimization methods yield limited gains, whereas post-training approaches (e.g., DPO) significantly enhance the model's contextual robustness. The results indicate that these issues do not stem from knowledge deficits in LLMs, but from a fundamental inability to maintain consistent reasoning under contextual distraction, posing a major challenge to the reliability of LLMs in real-world applications. The code is publicly available at https://github.com/wyf23187/Adaptive_Distractions . 39th Conference on Neural Information Processing Systems (NeurIPS 2025).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a78165b7-86fa-4957-beed-5471a782ecc6Cited by top-tier papers4
- ProbeLLM: Automating Principled Diagnosis of LLM FailuresYue Huang, Zhengzhe Jiang, Yuchen Ma, Yu Jiang et al.ICML 2026 · 4 citations
- PRISM-XR: Empowering Privacy-Aware XR Collaboration with Multimodal Large Language ModelsJiangong Chen, Mingyu Zhu, Bin LiIEEE VR 2026
- DynaSchedBench: Calibrated Dynamic Scheduling Benchmarks and Observability Paradox in LLM-based Scheduling AgentsShijie Cao, Yuan Yuan, Jing LiuICML 2026
- DataGen: Unified Synthetic Dataset Generation via Large Language ModelsYue Huang, Siyuan Wu, Chujie Gao, Dongping Chen et al.ICLR 2025
Builds on16
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Large Language Models Can Be Easily Distracted by Irrelevant ContextFreda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales et al.ICML 2023 · 970 citations
- Aligning AI With Shared Human ValuesDan Hendrycks, Collin Burns, Steven Basart, Andrew Critch et al.ICLR 2021 · 878 citations
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le et al.ICLR 2023 · 681 citations
Related papers
- How Is LLM Reasoning Distracted by Irrelevant Context? An Analysis Using a Controlled BenchmarkMinglai Yang, Ethan Huang, Liang Zhang, Mihai Surdeanu et al.EMNLP 2025 · 2 citations
- LLMs can be easily Confused by Instructional DistractionsYerin Hwang, Yongil Kim, Jahyun Koo, Taegwan Kang et al.ACL 2025
- Is Your (Reasoning) Multimodal Language Model Vulnerable Toward Distractions?Ming Liu, Hao Chen, Jindong Wang, Liwen Wang et al.AAAI 2026
- CaliDist: Calibrating Large Language Models via Behavioral Robustness to DistractionMohammad Anas Jawad, Cornelia CarageaICML 2026
- CodeCrash: Exposing LLM Fragility to Misleading Natural Language in Code ReasoningMan Ho Lam, Chaozheng Wang, Jen-Tse Huang, Michael R. LyuNeurIPS 2025 · 16 citations
