A Causal Framework to Quantify the Robustness of Mathematical Reasoning with Language Models
Alessandro Stolfo, Zhijing Jin, Kumar Shridhar, Bernhard Schölkopf, Mrinmaya Sachan
摘要
We have recently witnessed a number of impressive results on hard mathematical reasoning problems with language models. At the same time, the robustness of these models has also been called into question; recent works have shown that models can rely on shallow patterns in the problem description when generating a solution. Building on the idea of behavioral testing, we propose a novel framework, which pins down the causal effect of various factors in the input, e.g., the surface form of the problem text, the operands, and math operators on the output solution. By grounding the behavioral analysis in a causal graph describing an intuitive reasoning process, we study the behavior of language models in terms of robustness and sensitivity to direct interventions in the input space. We apply our framework on a test bed of math word problems. Our analysis shows that robustness does not appear to continuously improve as a function of size, but the GPT-3 Davinci models (175B) achieve a dramatic improvement in both robustness and sensitivity compared to all other GPT variants. 1 * Equal contribution. 1 Our code and data are available at https://github. com/alestolfo/causal-math . Kyle could fit n 1 =26 drawings on each page. If he has n 2 =11 pages, the number of drawings he can make is ___. Kyle could fit n 1 =2 drawings on each page. If he has n 2 =143 pages, the number of drawings he can make is ___.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper18
- Unveiling Causal Reasoning in Large Language Models: Reality or Mirage?Haoang Chi, He Li, Wenjing Yang, Feng Liu 等NeurIPS 2024 · 被引用 124 次
- CLadder: A Benchmark to Assess Causal Reasoning Capabilities of Language ModelsZhijing Jin, Yuen Chen, Felix Leeb, Luigi Gresele 等NeurIPS 2023 · 被引用 74 次
- Causal Prompting: Debiasing Large Language Model Prompting Based on Front-Door AdjustmentCongzhi Zhang, Linhai Zhang, Jialong Wu, Yulan He 等AAAI 2025 · 被引用 42 次
- NextQuill: Causal Preference Modeling for Enhancing LLM PersonalizationXiaoyan Zhao, Juntao You, Yang Zhang, Wenjie Wang 等ICLR 2026 · 被引用 38 次
- Do Language Models Exhibit the Same Cognitive Biases in Problem Solving as Human Learners?Andreas Opedal, Alessandro Stolfo, Haruki Shirakami, Ying Jiao 等ICML 2024 · 被引用 27 次
它引用的顶会 Paper13
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 被引用 3,415 次
- Investigating Gender Bias in Language Models Using Causal Mediation AnalysisJesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian 等NeurIPS 2020 · 被引用 851 次
相关 Paper
- DynaMath: A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language ModelsChengke Zou, Xingang Guo, Rui Yang, Junyu Zhang 等ICLR 2025
- MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard PerturbationsKaixuan Huang, Jiacheng Guo, Zihao Li, Xiang Ji 等ICML 2025
- GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem SolversQintong Li, Leyang Cui, Xueliang Zhao, Lingpeng Kong 等ACL 2024 · 被引用 9 次
- Decoupling Understanding from Reasoning via Problem Space Mapping for Small-Scale Model ReasoningLi Wang, Changhao Zhang, Zengqi Xiu, Kai Lu 等AAAI 2026 · 被引用 1 次
- Exposing the Achilles' Heel: Evaluating LLMs Ability to Handle Mistakes in Mathematical ReasoningJoykirat Singh, Akshay Uttama Nambi, Vibhav VineetACL 2025 · 被引用 10 次
