Can LLMs Reason Abstractly Over Math Word Problems Without CoT? Disentangling Abstract Formulation From Arithmetic Computation
Ziling Cheng, Meng Cao, Leila Pishdad, Yanshuai Cao, Jackie CK Cheung
Abstract
Final-answer-based metrics are commonly used for evaluating large language models (LLMs) on math word problems, often taken as proxies for reasoning ability. However, such metrics conflate two distinct sub-skills: abstract formulation (capturing mathematical relationships using expressions) and arithmetic computation (executing the calculations). Through a disentangled evaluation on GSM8K and SVAMP, we find that the final-answer accuracy of Llama-3 and Qwen2.5 (1B-32B) without CoT is overwhelmingly bottlenecked by the arithmetic computation step and not by the abstract formulation step. Contrary to the common belief, we show that CoT primarily aids in computation, with limited impact on abstract formulation. Mechanistically, we show that these two skills are composed conjunctively even in a single forward pass without any reasoning steps via an abstractthen-compute mechanism: models first capture problem abstractions, then handle computation. Causal patching confirms these abstractions are present, transferable, composable, and precede computation. These behavioural and mechanistic findings highlight the need for disentangled evaluation to accurately assess LLM reasoning and to guide future improvements. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 00627b2d-df49-4fa6-be25-8d23a33ce5a2Cited by top-tier papers3
- In-Context AlgebraEric Todd, Jannik Brinkmann, Rohit Gandikota, David BauICLR 2026 · 3 citations
- From Abstract to Contextual: What LLMs Still Cannot Do in MathematicsBowen Cao, Dongdong Zhang, Yixia Li, Junpeng Liu et al.ICLR 2026 · 1 citation
- Stochastic Chameleons: Irrelevant Context Hallucinations Reveal Class-Based (Mis)Generalization in LLMsZiling Cheng, Meng Cao, Marc-Antoine Rondeau, Jackie CK CheungACL 2025
Builds on17
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 3,415 citations
- Large Language Models Can Be Easily Distracted by Irrelevant ContextFreda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales et al.ICML 2023 · 970 citations
- PAL: Program-aided Language ModelsLuyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon et al.ICML 2023 · 700 citations
- Towards Best Practices of Activation Patching in Language Models: Metrics and MethodsFred Zhang, Neel NandaICLR 2024 · 233 citations
Related papers
- An Investigation of Neuron Activation as a Unified Lens to Explain Chain-of-Thought Eliciting Arithmetic Reasoning of LLMsDaking Rai, Ziyu YaoACL 2024
- CoMAT: Chain of Mathematically Annotated Thought Improves Mathematical ReasoningJoshua Ong Jun Leang, Aryo Pradipta Gema, Shay B. CohenEMNLP 2025
- Benchmarking Abstract and Reasoning Abilities Through A Theoretical PerspectiveQingchuan Ma, Yuhang Wu, Xiawu Zheng, Rongrong JiICML 2025
- To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoningZayne Rea Sprague, Fangcong Yin, Juan Diego Rodriguez, Dongwei Jiang et al.ICLR 2025
- Understanding Chain-of-Thought in LLMs through Information TheoryJean-Francois Ton, Muhammad Faaiz Taufiq, Yang LiuICML 2025
