Case-Based or Rule-Based: How Do Transformers Do the Math?
Yi Hu, Xiaojuan Tang, Haotong Yang, Muhan Zhang
摘要
Despite the impressive performance in a variety of complex tasks, modern large language models (LLMs) still have trouble dealing with some math problems that are simple and intuitive for humans, such as addition. While we can easily learn basic rules of addition and apply them to new problems of any length, LLMs struggle to do the same. Instead, they may rely on similar cases seen in the training corpus for help. We define these two different reasoning mechanisms as "rule-based reasoning" and "case-based reasoning". Since rule-based reasoning is essential for acquiring systematic generalization ability, we aim to explore exactly whether transformers use rule-based or case-based reasoning for math problems. Through carefully designed intervention experiments on five math tasks, we confirm that transformers are performing case-based reasoning, no matter whether scratchpad is used, which aligns with the previous observations that transformers use subgraph matching/shortcut learning to reason. To mitigate such problems, we propose a Rule-Following Fine-Tuning (RFFT) technique to teach transformers to perform rule-based reasoning. Specifically, we provide explicit rules in the input and then instruct transformers to recite and follow the rules step by step. Through RFFT, we successfully enable LLMs fine-tuned on 1-5 digit addition to generalize to up to 12digit addition with over 95% accuracy, which is over 40% higher than scratchpad. The significant improvement demonstrates that teaching LLMs to use rules explicitly helps them learn rulebased reasoning and generalize better in length.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper19
- Embedding Trajectory for Out-of-Distribution Detection in Mathematical ReasoningYiming Wang, Pei Zhang, Baosong Yang, Derek F. Wong 等NeurIPS 2024 · 被引用 23 次
- Mixture-of-Experts Can Surpass Dense LLMs Under Strictly Equal ResourceHouyi Li, Ka Man Lo, Shijie Xuyang, Ziqi Wang 等ICLR 2026 · 被引用 8 次
- Beyond Single-Task: Robust Multi-Task Length Generalization for LLMsYi Hu, Shijia Kang, Haotong Yang, Haotian Xu 等NeurIPS 2025 · 被引用 6 次
- Physically Ground Commonsense Knowledge for Articulated Object Manipulation with Analytic ConceptsJiude Wei, Yuxuan Li, Cewu Lu, Jianhua SunCVPR 2026 · 被引用 6 次
- Curing Miracle Steps in LLM Mathematical Reasoning with Rubric RewardsYouliang Yuan, Qiuyang Mang, Jingbang Chen, Hong Wan 等ACL 2026 · 被引用 5 次
它引用的顶会 Paper21
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- What Can Transformers Learn In-Context? A Case Study of Simple Function ClassesShivam Garg, Dimitris Tsipras, Percy Liang, Gregory ValiantNeurIPS 2022 · 被引用 883 次
- Are Transformers universal approximators of sequence-to-sequence functions?Chulhee Yun, Srinadh Bhojanapalli, Ankit Singh Rawat, Sashank J. Reddi 等ICLR 2020 · 被引用 481 次
相关 Paper
- Exploring Length Generalization in Large Language ModelsCem Anil, Yuhuai Wu, Anders Andreassen, Aitor Lewkowycz 等NeurIPS 2022 · 被引用 267 次
- Incentivizing LLM Reasoning via Reinforcement Learning with Functional Monte Carlo Tree SearchKongcheng Zhang, QI YAO, Baisheng Lai, Jiaxing Huang 等ICLR 2026
- The Emperor's New Reasoning: Format Imitation Overshadows Genuine Mathematical Understanding in SFTLinyao Yang, Jian-Tao Huang, Yafei Lu, Zhenhui Jessie Li 等EMNLP 2025
- Do Large Language Models Truly Grasp Addition? A Rule-Focused Diagnostic Using Two-Integer ArithmeticYang Yan, Yu Lu, Renjun Xu, Zhenzhong LanEMNLP 2025 · 被引用 1 次
- ReFT: Reasoning with Reinforced Fine-TuningLuong Quoc Trung, Xinbo Zhang, Zhanming Jie, Peng Sun 等ACL 2024
