GSM-∞: How Do your LLMs Behave over Infinitely Increasing Reasoning Complexity and Context Length?
Yang Zhou, Hongyi Liu, Zhuoming Chen, Yuandong Tian, Beidi Chen
Abstract
Long-context large language models (LLMs) have recently shown strong performance in information retrieval and long-document QA. However, to tackle the most challenging intellectual problems, LLMs must reason effectively in long and complex contexts (e.g., frontier mathematical research). Studying how LLMs handle increasing reasoning complexity and context length is essential, yet existing benchmarks lack a solid basis for quantitative evaluation. Inspired by the abstraction of GSM-8K problems as computational graphs-and the ability to introduce noise by adding unnecessary nodes and edges-we develop a grade-school math problem generator capable of producing arithmetic problems with infinite difficulty and context length under finegrained control. Using our newly synthesized GSM-∞ benchmark, we comprehensively evaluate existing LLMs. We find a consistent sigmoid decline in reasoning performance as problem complexity increases, along with a systematic inference scaling trend: exponentially increasing inference computation yields only linear performance gains. These findings underscore the fundamental limitations of current longcontext LLMs and the key challenges in scaling reasoning capabilities. Our GSM-∞ benchmark provides a scalable and controllable testbed for systematically studying and advancing LLM reasoning in long and complex contexts. Code open-sources at https://infini-ai-lab. github.io/gsm_infinite/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0c67da05-a3e0-49b3-ba21-b78641e72c17Builds on8
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Active Retrieval Augmented GenerationZhengbao Jiang, Frank F. Xu, Luyu Gao, Zhiqing Sun et al.EMNLP 2023 · 315 citations
- LongBench: A Bilingual, Multitask Benchmark for Long Context UnderstandingYushi Bai, Xin Lv, Jiajie Zhang, Hongchang Lyu et al.ACL 2024 · 94 citations
- Same Task, More Tokens: the Impact of Input Length on the Reasoning Performance of Large Language ModelsMosh Levy, Alon Jacoby, Yoav GoldbergACL 2024 · 77 citations
- Physics of Language Models: Part 2.1, Grade-School Math and the Hidden Reasoning ProcessTian Ye, Zicheng Xu, Yuanzhi Li, Zeyuan Allen-ZhuICLR 2025 · 3 citations
Related papers
- ınftyBench: Extending Long Context Evaluation Beyond 100K TokensXinrong Zhang, Yingfa Chen, Shengding Hu, Zihang Xu et al.ACL 2024
- Language models are multilingual chain-of-thought reasonersFreda Shi, Mirac Suzgun, Markus Freitag, Xuezhi Wang et al.ICLR 2023 · 52 citations
- GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language ModelsIman Mirzadeh, Keivan Alizadeh, Hooman Shahrokhi, Oncel Tuzel et al.ICLR 2025
- Can LLMs Solve Longer Math Word Problems Better?Xin Xu, Tong Xiao, Zitong Chao, Zhenya Huang et al.ICLR 2025
- How Is LLM Reasoning Distracted by Irrelevant Context? An Analysis Using a Controlled BenchmarkMinglai Yang, Ethan Huang, Liang Zhang, Mihai Surdeanu et al.EMNLP 2025 · 2 citations
