Chain of Thoughtlessness? An Analysis of CoT in Planning
Kaya Stechly, Karthik Valmeekam, Subbarao Kambhampati
Abstract
Large language model (LLM) performance on reasoning problems typically does not generalize out of distribution. Previous work has claimed that this can be mitigated with chain of thought prompting-a method of demonstrating solution procedures-with the intuition that it is possible to in-context teach an LLM an algorithm for solving the problem. This paper presents a case study of chain of thought on problems from Blocksworld, a classical planning domain, and examines the performance of two state-of-the-art LLMs across two axes: generality of examples given in prompt, and complexity of problems queried with each prompt. While our problems are very simple, we only find meaningful performance improvements from chain of thought prompts when those prompts are exceedingly specific to their problem class, and that those improvements quickly deteriorate as the size n of the query-specified stack grows past the size of stacks shown in the examples. We also create scalable variants of three domains commonly studied in previous CoT papers and demonstrate the existence of similar failure modes. Our results hint that, contrary to previous claims in the literature, CoT's performance improvements do not stem from the model learning general algorithmic procedures via demonstrations but depend on carefully engineering highly problem specific prompts. This spotlights drawbacks of chain of thought, especially the sharp tradeoff between possible performance gains and the amount of human labor necessary to generate examples with correct reasoning traces.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 52a9538c-c3a9-456b-a8d1-6f441dd4dfd1Cited by top-tier papers32
- Curriculum Reinforcement Learning from Easy to Hard Tasks Improves LLM ReasoningShubham Parashar, Shurui Gui, Xiner Li, Hongyi Ling et al.ICLR 2026 · 112 citations
- The Illusion of Diminishing Returns: Measuring Long Horizon Execution in LLMsAkshit Sinha, Arvindh Arun, Shashwat Goel, Steffen Staab et al.ICLR 2026 · 64 citations
- Long-Horizon Planning for Multi-Agent Robots in Partially Observable EnvironmentsSiddharth Nayak, Adelmo Morrison Orozco, Marina Ten Have, Jackson Zhang et al.NeurIPS 2024 · 41 citations
- FaithCoT-Bench: Benchmarking Instance-Level Faithfulness of Chain-of-Thought ReasoningXu Shen, Song Wang, Zhen Tan, Laura Yao et al.ICLR 2026 · 28 citations
- Classical Planning with LLM-Generated Heuristics: Challenging the State of the Art with Python CodeAugusto B. Corrêa, André Grahl Pereira, Jendrik SeippNeurIPS 2025 · 27 citations
Builds on23
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought PromptingMiles Turpin, Julian Michael, Ethan Perez, Samuel R. BowmanNeurIPS 2023 · 1,792 citations
- Improving Factuality and Reasoning in Language Models through Multiagent DebateYilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum et al.ICML 2024 · 1,562 citations
- Why Johnny Can't Prompt: How Non-AI Experts Try (and Fail) to Design LLM PromptsJ. D. Zamfirescu-Pereira, Richmond Y. Wong, Bjoern Hartmann, Qian YangCHI 2023 · 892 citations
Related papers
- Why Prompt Design Matters and Works: A Complexity Analysis of Prompt Search Space in LLMsXiang Zhang, Juntai Cao, Chenyu You, Dujian DingACL 2025 · 21 citations
- Automatic Chain of Thought Prompting in Large Language ModelsZhuosheng Zhang, Aston Zhang, Mu Li, Alex SmolaICLR 2023 · 234 citations
- Unveiling Factual Recall Behaviors of Large Language Models through Knowledge NeuronsYifei Wang, Yuheng Chen, Wanting Wen, Yu Sheng et al.EMNLP 2024 · 3 citations
- To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoningZayne Rea Sprague, Fangcong Yin, Juan Diego Rodriguez, Dongwei Jiang et al.ICLR 2025
- Towards Revealing the Mystery behind Chain of Thought: A Theoretical PerspectiveGuhao Feng, Bohang Zhang, Yuntian Gu, Haotian Ye et al.NeurIPS 2023 · 470 citations
