Code Prompting Elicits Conditional Reasoning Abilities in Text+Code LLMs
Haritz Puerto, Martin Tutek, Somak Aditya, Xiaodan Zhu, Iryna Gurevych
摘要
Reasoning is a fundamental component of language understanding. Recent prompting techniques, such as chain of thought, have consistently improved LLMs' performance on various reasoning tasks. Nevertheless, there is still little understanding of what triggers reasoning abilities in LLMs in the inference stage. In this paper, we investigate the effect of the input representation on the reasoning abilities of LLMs. We hypothesize that representing natural language tasks as code can enhance specific reasoning abilities such as entity tracking or logical reasoning. To study this, we propose code prompting, a methodology we operationalize as a chain of prompts that transforms a natural language problem into code and directly prompts the LLM using the generated code without resorting to external code execution. We find that code prompting exhibits a high-performance boost for multiple LLMs (up to 22.52 percentage points on GPT 3.5, 7.75 on Mixtral, and 16.78 on Mistral) across multiple conditional reasoning datasets. We then conduct comprehensive experiments to understand how the code representation triggers reasoning abilities and which capabilities are elicited in the underlying models. Our analysis on GPT 3.5 reveals that the code formatting of the input problem is essential for performance improvement. Furthermore, the code representation improves sample efficiency of in-context learning and facilitates state tracking of entities. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Synatra: Turning Indirect Knowledge into Direct Demonstrations for Digital Agents at ScaleTianyue Ou, Frank F. Xu, Aman Madaan, Jiarui Liu 等NeurIPS 2024 · 被引用 45 次
- Fine-Tuning on Diverse Reasoning Chains Drives Within-Inference CoT Refinement in LLMsHaritz Puerto, Tilek Chubakov, Xiaodan Zhu, Harish Tayyar Madabushi 等ACL 2025 · 被引用 13 次
- Read it in Two Steps: Translating Extremely Low-Resource Languages with Code-Augmented Grammar BooksChen Zhang, Jiuheng Lin, Xiao Liu, Zekai Zhang 等ACL 2025 · 被引用 5 次
- From Tools to Teammates: Evaluating LLMs in Multi-Session Coding InteractionsNathanaël Carraz Rakotonirina, Mohammed Hamdy, Jon Ander Campos, Lucas Weber 等ACL 2025 · 被引用 5 次
- Code to Think, Think to Code: A Survey on Code-Enhanced Reasoning and Reasoning-Driven Code Intelligence in LLMsDayu Yang, Tianyang Liu, Daoan Zhang, Antoine Simoulin 等EMNLP 2025 · 被引用 1 次
它引用的顶会 Paper14
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- PAL: Program-aided Language ModelsLuyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon 等ICML 2023 · 被引用 700 次
- FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text GenerationSewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis 等EMNLP 2023 · 被引用 225 次
相关 Paper
- Chain of Code: Reasoning with a Language Model-Augmented Code EmulatorChengshu Li, Jacky Liang, Andy Zeng, Xinyun Chen 等ICML 2024 · 被引用 155 次
- CodePlan: Unlocking Reasoning Potential in Large Language Models by Scaling Code-form PlanningJiaxin Wen, Jian Guan, Hongning Wang, Wei Wu 等ICLR 2025
- NL-Debugging: Exploiting Natural Language as an Intermediate Representation for Code DebuggingWeiming Zhang, Qingyao Li, Xinyi Dai, Jizheng Chen 等EMNLP 2025 · 被引用 1 次
- At Which Training Stage Does Code Data Help LLMs Reasoning?Yingwei Ma, Yue Liu, Yue Yu, Yuanliang Zhang 等ICLR 2024 · 被引用 106 次
- CodeIE: Large Code Generation Models are Better Few-Shot Information ExtractorsPeng Li, Tianxiang Sun, Qiong Tang, Hang Yan 等ACL 2023 · 被引用 41 次
