From A and B to A+B: Can Large Language Models Solve Compositional Math Problems?
Xisheng Xiao, Hanlin Zhao
摘要
Large language models (LLMs) have demonstrated strong performance in solving math problems, and there is growing research on evaluating their robustness. Unlike previous studies that create problem variants by adding perturbations to a single problem, this paper focuses on the interaction between problems. Specifically, we combine two original problems with a logical connection to get a new math problem, and measure the LLM's performance on it to evaluate its compositional generalization, which is an important and essential reasoning capability in human intelligence. The result of experiments that cover 14 different LLMs shows that even when the mathematical essence remains unchanged, a simple form of combination can significantly reduce the performance of LLMs, revealing the limitation of their generalization ability. Additionally, we propose an automated pipeline with 98.2% accuracy to assist in annotating datasets (1 manual, 2 synthetic). The extensive experiments conducted on these datasets further verify the conclusion and obtain some important findings. Finally, we analyze the impact of factors such as difficulty and length on LLMs' performance, offering insights for future research. Code and data are avaliable at https://github.com/goooooodcoder/NCSP .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- Large Language Models Can Be Easily Distracted by Irrelevant ContextFreda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales 等ICML 2023 · 被引用 970 次
- MetaMath: Bootstrap Your Own Mathematical Questions for Large Language ModelsLonghui Yu, Weisen Jiang, Han Shi, Jincheng Yu 等ICLR 2024 · 被引用 637 次
- The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A"Lukas Berglund, Meg Tong, Maximilian Kaufmann, Mikita Balesni 等ICLR 2024 · 被引用 462 次
- ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem SolvingZhibin Gou, Zhihong Shao, Yeyun Gong, Yelong Shen 等ICLR 2024 · 被引用 289 次
相关 Paper
- MathGAP: Out-of-Distribution Evaluation on Problems with Arbitrarily Complex ProofsAndreas Opedal, Haruki Shirakami, Bernhard Schölkopf, Abulhair Saparov 等ICLR 2025
- GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem SolversQintong Li, Leyang Cui, Xueliang Zhao, Lingpeng Kong 等ACL 2024 · 被引用 9 次
- Faith and Fate: Limits of Transformers on CompositionalityNouha Dziri, Ximing Lu, Melanie Sclar, Xiang Lorraine Li 等NeurIPS 2023 · 被引用 728 次
- AgentCoMa: A Compositional Benchmark Mixing Commonsense and Mathematical Reasoning in Real-World ScenariosLisa Alazraki, Lihu Chen, Ana Brassard, Joe Stacey 等ACL 2026 · 被引用 1 次
- Can LLMs Solve Longer Math Word Problems Better?Xin Xu, Tong Xiao, Zitong Chao, Zhenya Huang 等ICLR 2025
