CodeCrash: Exposing LLM Fragility to Misleading Natural Language in Code Reasoning
Man Ho Lam, Chaozheng Wang, Jen-Tse Huang, Michael R. Lyu
Abstract
Large Language Models (LLMs) have recently demonstrated strong capabilities in code-related tasks, but their robustness in code reasoning under perturbations remains underexplored. We introduce CODECRASH, a stress-testing framework with 1,279 questions from CRUXEVAL and LIVECODEBENCH, designed to evaluate reasoning reliability under structural perturbations and misleading natural language (NL) contexts. Through a systematic evaluation of 17 LLMs, we find that models often shortcut reasoning by over-relying on NL cues, leading to an average performance degradation of 23.2% in output prediction tasks. Even with Chain-of-Thought reasoning, models on average still have a 13.8% drop due to distractibility and rationalization, revealing a lack of critical reasoning capability to distinguish the actual code behaviors. While Large Reasoning Models with internal reasoning mechanisms improve robustness by fostering critical thinking, plausible yet incorrect hints can trigger pathological self-reflection, causing 2 -3 times token consumption and even catastrophic cognitive dissonance in extreme cases for QwQ-32B. We refer to this phenomenon as Reasoning Collapse. CODECRASH provides a rigorous benchmark for evaluating robustness in code reasoning, guiding future research and development toward more reliable and resilient models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3c1b65af-c0b5-47a8-a81c-f85d8d834768Cited by top-tier papers4
- Verify Before You Commit: Towards Faithful Reasoning in LLM Agents via Self-AuditingWenhao Yuan, Chenchen Lin, Jian Chen, Jinfeng Xu et al.ACL 2026 · 3 citations
- The Deception Delta: Adversarial Evaluation of LLM-Based Smart Contract Bytecode ForensicsTimo SchefoldCCS 2026
- The Invisible Lottery: How Subtle Cues Steer Algorithm Choice in LLM Code GenerationAkanksha Narula, Mofasshara Rafique, Laurent BindschaedlerICML 2026
- CREME: Robustness Enhancement of Code LLMs via Layer-Aware Model EditingShuhan Liu, Xing Hu, Kerui Huang, Xiaohu Yang et al.ICSE 2026
Builds on20
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 2,317 citations
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao et al.ICLR 2024 · 2,082 citations
- Large Language Models Can Be Easily Distracted by Irrelevant ContextFreda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales et al.ICML 2023 · 970 citations
- LEVER: Learning to Verify Language-to-Code Generation with ExecutionAnsong Ni, Srini Iyer, Dragomir Radev, Veselin Stoyanov et al.ICML 2023 · 318 citations
Related papers
- Reasoning Runtime Behavior of a Program with LLM: How Far are We?Junkai Chen, Zhiyuan Pan, Xing Hu, Zhenhao Li et al.ICSE 2025 · 5 citations
- UniCode: Augmenting Evaluation for Code ReasoningXinyue Zheng, Haowei Lin, Shaofei Cai, Yaodong Yang et al.ICML 2026 · 1 citation
- CRUXEVAL-X: A Benchmark for Multilingual Code Reasoning, Understanding and ExecutionRuiyang Xu, Jialun Cao, Yaojie Lu, Ming Wen et al.ACL 2025 · 26 citations
- DARG: Dynamic Evaluation of Large Language Models via Adaptive Reasoning GraphZhehao Zhang, Jiaao Chen, Diyi YangNeurIPS 2024 · 42 citations
- CRUXEval: A Benchmark for Code Reasoning, Understanding and ExecutionAlex Gu, Baptiste Rozière, Hugh James Leather, Armando Solar-Lezama et al.ICML 2024 · 270 citations
