Reasoning Robustness of LLMs to Adversarial Typographical Errors
Esther Gan, Yiran Zhao, Liying Cheng, Yancan Mao, Anirudh Goyal, Kenji Kawaguchi, Min-Yen Kan, Michael Shieh
Abstract
Large Language Models (LLMs) have demonstrated impressive capabilities in reasoning using Chain-of-Thought (CoT) prompting. However, CoT can be biased by users' instruction. In this work, we study the reasoning robustness of LLMs to typographical errors, which can naturally occur in users' queries. We design an Adversarial Typo Attack (ATA) algorithm that iteratively samples typos for words that are important to the query and selects the edit that is most likely to succeed in attacking. It shows that LLMs are sensitive to minimal adversarial typographical changes. Notably, with 1 character edit, Mistral-7B-Instruct's accuracy drops from 43.7% to 38.6% on GSM8K, while with 8 character edits the performance further drops to 19.2%. To extend our evaluation to larger and closed-source LLMs, we develop the R 2 ATA benchmark, which assesses models' Reasoning Robustness to ATA. It includes adversarial typographical questions derived from three widely-used reasoning datasets-GSM8K, BBH, and MMLU-by applying ATA to open-source LLMs. R 2 ATA demonstrates remarkable transferability and causes notable performance drops across multiple super large and closed-source LLMs. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5f937020-3bfd-4b5b-b29b-dccfdc8e4ae1Cited by top-tier papers10
- One Token Embedding Is Enough to Deadlock Your Large Reasoning ModelMohan Zhang, Yihua Zhang, Jinghan Jia, Zhangyang (Atlas) Wang et al.NeurIPS 2025 · 7 citations
- Evaluating Robustness of Large Language Models Against Multilingual Typographical ErrorsRaoyuan Zhao, Yihong Liu, Lena Altinger, Hinrich Schütze et al.ACL 2026 · 6 citations
- Bounds of Chain-of-Thought Robustness: Reasoning Steps, Embed Norms, and BeyondDingzirui Wang, Xuanliang Zhang, Keyan Xu, Qingfu Zhu et al.ICLR 2026 · 3 citations
- BadThink: Triggered Overthinking Attacks on Chain-of-Thought Reasoning in Large Language ModelsShuaitong Liu, Renjue Li, Lijia Yu, Lijun Zhang et al.AAAI 2026
- Mastering Board Games by External and Internal Planning with Language ModelsJohn Schultz, Jakub Adámek, Matej Jusup, Marc Lanctot et al.ICML 2025
Builds on7
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought PromptingMiles Turpin, Julian Michael, Ethan Perez, Samuel R. BowmanNeurIPS 2023 · 1,792 citations
- Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and EntailmentDi Jin, Zhijing Jin, Joey Tianyi Zhou, Peter SzolovitsAAAI 2020 · 1,333 citations
- Large Language Models Can Be Easily Distracted by Irrelevant ContextFreda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales et al.ICML 2023 · 970 citations
Related papers
- GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem SolversQintong Li, Leyang Cui, Xueliang Zhao, Lingpeng Kong et al.ACL 2024 · 9 citations
- MathAttack: Attacking Large Language Models towards Math Solving AbilityZihao Zhou, Qiufeng Wang, Mingyu Jin, Jie Yao et al.AAAI 2024 · 38 citations
- Can LLMs Learn from Previous Mistakes? Investigating LLMs' Errors to Boost for ReasoningYongqi Tong, Dawei Li, Sizhe Wang, Yujia Wang et al.ACL 2024
- REALISTA: Realistic Latent Adversarial Attacks that Elicit LLM HallucinationsBuyun Liang, Jinqi Luo, Liangzu Peng, Kwan Ho Ryan Chan et al.ICML 2026
- mCoT: Multilingual Instruction Tuning for Reasoning Consistency in Language ModelsHuiyuan Lai, Malvina NissimACL 2024 · 5 citations
