LongCoT: Benchmarking Long-Horizon Chain-of-Thought Reasoning
Sumeet Motwani, Daniel Nichols, Charles London, Peggy Li, Fabio Pizzati, Acer Blake, Hasan Hammoud, Tavish McDonald, Akshat Naik, Alesia Ivanova, Vignesh Baskaran, Ivan Laptev
Abstract
As language models are increasingly deployed for complex autonomous tasks, their ability to reason accurately over longer horizons becomes critical. An essential component of this ability is planning and managing a long, complex chain-of-thought (CoT). We introduce LongCoT, a scalable benchmark of 2,500 expert-designed problems spanning chemistry, mathematics, computer science, chess, and logic to isolate and directly measure the long-horizon CoT reasoning capabilities of frontier models. Problems consist of a short input with a verifiable answer; solving them requires navigating a graph of interdependent steps that span tens to hundreds of thousands of reasoning tokens. Each local step is individually tractable for frontier models, so failures reflect long-horizon reasoning limitations. At release, the best models achieve <10% accuracy (GPT 5.2: 9.8%; Gemini 3 Pro: 6.1%) on LongCoT, revealing a substantial gap in current capabilities. Overall, LongCoT provides a rigorous measure of long-horizon reasoning, tracking the ability of frontier models to reason reliably over extended periods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 39ed1197-5008-41c2-b2e2-a691d5391303Builds on7
- Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line InterfacesMike A. Merrill, Alexander Glenn Shaw, Nicholas Carlini, Boxuan Li et al.ICLR 2026 · 520 citations
- The Illusion of Diminishing Returns: Measuring Long Horizon Execution in LLMsAkshit Sinha, Arvindh Arun, Shashwat Goel, Steffen Staab et al.ICLR 2026 · 64 citations
- RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable EnvironmentsZhiyuan Zeng, Hamish Ivison, Yiping Wang, Lifan Yuan et al.ICML 2026 · 28 citations
- h1: Bootstrapping LLMs to Reason over Longer Horizons via Reinforcement LearningAlesia Ivanova, Sumeet Motwani, Jack Cai, Phil Torr et al.ICML 2026 · 11 citations
- Generating Creative Chess PuzzlesXidong Feng, Vivek Veeriah, Marcus Chiam, Michael Dennis et al.NeurIPS 2025 · 7 citations
Related papers
- VideoReasonBench: Can MLLMs Perform Vision-Centric Complex Video Reasoning?Yuanxin Liu, Kun Ouyang, Haoning Wu, Yi Liu et al.ICLR 2026 · 20 citations
- Reasoning about Reasoning: BAPO Bounds on Chain-of-Thought Token Complexity in LLMsKiran Tomlinson, Tobias Schnabel, Adith Swaminathan, Jennifer NevilleICML 2026 · 2 citations
- R-Horizon: How Far Can Your Large Reasoning Model Really Go in Breadth and Depth?Yi Lu, Jianing Wang, Linsen Guo, Wei He et al.ICLR 2026 · 9 citations
- Lost in Transmission: When and Why LLMs Fail to Reason GloballyTobias Schnabel, Kiran Tomlinson, Adith Swaminathan, Jennifer NevilleNeurIPS 2025 · 9 citations
- RefineBench: Evaluating Refinement Capability of Language Models via ChecklistsYoung-Jun Lee, Seungone Kim, Byung-Kwan Lee, Minkyeong Moon et al.ICLR 2026 · 13 citations
