On Code-Induced Reasoning in LLMs
Abdul Waheed, Zhen Wu, Carolyn Rose, Daphne Ippolito
摘要
Code data has been shown to enhance the reasoning capabilities of large language models (LLMs), but it remains unclear which aspects of code are most responsible. We investigate this question with a systematic, data-centric framework. We construct parallel instruction datasets across ten programming languages and introduce controlled perturbations that selectively disrupt structural and semantic properties of code. We then fine-tune LLMs from five model families and eight scales on each variant and evaluate their performance on natural language, math, and code tasks. Across 3,331 experiments, our results show that LLMs are more vulnerable to structural perturbations than semantic ones, particularly on math and code tasks. Appropriate abstractions like pseudocode and flowcharts can be as effective as code, while encoding the same information with fewer tokens without adhering to original syntax can often retain or even improve performance. Notably, even corrupted code with misleading signals remains competitive when surface-level regularities persist. Finally, syntactic styles also shape task-specific gains, with Python favoring natural language reasoning and lower-level languages such as Java and Rust favoring math. Through our systematic framework, we provide a fine-grained analysis of how different aspects of code influence reasoning and inform the design of training data for enhancing LLM reasoning capabilities.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Deduplicating Training Data Makes Language Models BetterKatherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang 等ACL 2022 · 被引用 844 次
- ZeRO-Offload: Democratizing Billion-Scale Model TrainingJie Ren, Samyam Rajbhandari, Reza Yazdani Aminabadi, Olatunji Ruwase 等USENIX ATC 2021 · 被引用 657 次
- At Which Training Stage Does Code Data Help LLMs Reasoning?Yingwei Ma, Yue Liu, Yue Yu, Yuanliang Zhang 等ICLR 2024 · 被引用 106 次
相关 Paper
- When Do Program-of-Thought Works for Reasoning?Zhen Bi, Ningyu Zhang, Yinuo Jiang, Shumin Deng 等AAAI 2024
- To Code or Not To Code? Exploring Impact of Code in Pre-trainingViraat Aryabumi, Yixuan Su, Raymond Ma, Adrien Morisot 等ICLR 2025 · 被引用 3 次
- Unveiling the Impact of Coding Data Instruction Fine-Tuning on Large Language Models ReasoningXinlu Zhang, Zhiyu Zoey Chen, Xi Ye, Xianjun Yang 等AAAI 2025 · 被引用 40 次
- Procedural Pretraining: Warming Up Language Models with Abstract DataLiangze Jiang, Zachary Shinnick, Anton Hengel, Hemanth Saratchandran 等ICML 2026 · 被引用 6 次
- What Really Improves Mathematical Reasoning: Structured Reasoning Signals Beyond Pure CodeYuze Zhao, Junpeng Fang, Lu Yu, Zhenya Huang 等ICML 2026 · 被引用 1 次
