Reasoning Robustness of LLMs to Adversarial Typographical Errors
Esther Gan, Yiran Zhao, Liying Cheng, Yancan Mao, Anirudh Goyal, Kenji Kawaguchi, Min-Yen Kan, Michael Shieh
摘要
Large Language Models (LLMs) have demonstrated impressive capabilities in reasoning using Chain-of-Thought (CoT) prompting. However, CoT can be biased by users' instruction. In this work, we study the reasoning robustness of LLMs to typographical errors, which can naturally occur in users' queries. We design an Adversarial Typo Attack (ATA) algorithm that iteratively samples typos for words that are important to the query and selects the edit that is most likely to succeed in attacking. It shows that LLMs are sensitive to minimal adversarial typographical changes. Notably, with 1 character edit, Mistral-7B-Instruct's accuracy drops from 43.7% to 38.6% on GSM8K, while with 8 character edits the performance further drops to 19.2%. To extend our evaluation to larger and closed-source LLMs, we develop the R 2 ATA benchmark, which assesses models' Reasoning Robustness to ATA. It includes adversarial typographical questions derived from three widely-used reasoning datasets-GSM8K, BBH, and MMLU-by applying ATA to open-source LLMs. R 2 ATA demonstrates remarkable transferability and causes notable performance drops across multiple super large and closed-source LLMs. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- One Token Embedding Is Enough to Deadlock Your Large Reasoning ModelMohan Zhang, Yihua Zhang, Jinghan Jia, Zhangyang (Atlas) Wang 等NeurIPS 2025 · 被引用 7 次
- Evaluating Robustness of Large Language Models Against Multilingual Typographical ErrorsRaoyuan Zhao, Yihong Liu, Lena Altinger, Hinrich Schütze 等ACL 2026 · 被引用 6 次
- Bounds of Chain-of-Thought Robustness: Reasoning Steps, Embed Norms, and BeyondDingzirui Wang, Xuanliang Zhang, Keyan Xu, Qingfu Zhu 等ICLR 2026 · 被引用 3 次
- BadThink: Triggered Overthinking Attacks on Chain-of-Thought Reasoning in Large Language ModelsShuaitong Liu, Renjue Li, Lijia Yu, Lijun Zhang 等AAAI 2026
- Mastering Board Games by External and Internal Planning with Language ModelsJohn Schultz, Jakub Adámek, Matej Jusup, Marc Lanctot 等ICML 2025
它引用的顶会 Paper7
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought PromptingMiles Turpin, Julian Michael, Ethan Perez, Samuel R. BowmanNeurIPS 2023 · 被引用 1,792 次
- Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and EntailmentDi Jin, Zhijing Jin, Joey Tianyi Zhou, Peter SzolovitsAAAI 2020 · 被引用 1,333 次
- Large Language Models Can Be Easily Distracted by Irrelevant ContextFreda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales 等ICML 2023 · 被引用 970 次
相关 Paper
- GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem SolversQintong Li, Leyang Cui, Xueliang Zhao, Lingpeng Kong 等ACL 2024 · 被引用 9 次
- MathAttack: Attacking Large Language Models towards Math Solving AbilityZihao Zhou, Qiufeng Wang, Mingyu Jin, Jie Yao 等AAAI 2024 · 被引用 38 次
- Can LLMs Learn from Previous Mistakes? Investigating LLMs' Errors to Boost for ReasoningYongqi Tong, Dawei Li, Sizhe Wang, Yujia Wang 等ACL 2024
- REALISTA: Realistic Latent Adversarial Attacks that Elicit LLM HallucinationsBuyun Liang, Jinqi Luo, Liangzu Peng, Kwan Ho Ryan Chan 等ICML 2026
- mCoT: Multilingual Instruction Tuning for Reasoning Consistency in Language ModelsHuiyuan Lai, Malvina NissimACL 2024 · 被引用 5 次
