How Does Chain of Thought Think? Mechanistic Interpretability of Chain-of-Thought Reasoning with Sparse Autoencoding
Xi Chen, Aske Plaat, Niki van Stein
Abstract
Chain‑of‑thought (CoT) prompting boosts Large Language Models accuracy on multi‑step tasks, yet whether the generated ``thoughts'' reflect the true internal reasoning process is unresolved. We present the first feature‑level causal study of CoT faithfulness. Combining sparse autoencoders with activation patching, we extract monosemantic features from Pythia‑70M and Pythia‑2.8B while they tackle GSM8K math problems under CoT and plain (noCoT) prompting. Swapping a small set of CoT‑reasoning features into a noCoT run raises answer log‑probabilities significantly in the 2.8B model, but has no reliable effect in 70M, revealing a clear contrast for these two scales. CoT also leads to significantly higher activation sparsity and feature interpretability scores in the larger model, signalling more modular internal computation. For example, the model's confidence in generating correct answers improves from 1.2 to 4.3. We introduce patch‑curves and random‑feature patching baselines, showing that useful CoT information is not only present in the top-K patches but widely distributed. Overall, our results indicate that CoT can induce more interpretable internal structures in high-capacity LLMs, validating its role as a structured prompting method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dc652d6c-534f-49d0-a741-591826d3435fCited by top-tier papers4
- Do Sparse Autoencoders Identify Reasoning Features in Language Models?George Ma, Zhongyuan Liang, Irene Y. Chen, Somayeh SojoudiICML 2026 · 9 citations
- Good Arguments Against the People Pleasers: How Reasoning Mitigates (Yet Masks) LLM SycophancyZhaoxin Feng, Zheng Chen, Jianfei Ma, Yip Tin Po et al.ACL 2026 · 3 citations
- What "Not" to Detect: Negation-Aware VLMs via Structured Reasoning and Token MergingInha Kang, Youngsun Lim, Seonho Lee, Jiho Choi et al.ICLR 2026 · 1 citation
- The Tell-Tale Norm: Magnitude as a Signal for Reasoning Dynamics in Large Language ModelsJinyang Zhang, Hongxin Ding, Yue Fang, Weibin Liao et al.ICML 2026
Builds on16
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 3,415 citations
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart et al.ICLR 2024 · 1,072 citations
- Causal Abstractions of Neural NetworksAtticus Geiger, Hanson Lu, Thomas Icard, Christopher PottsNeurIPS 2021 · 516 citations
Related papers
- Chain-of-Thought Reasoning Without PromptingXuezhi Wang, Denny ZhouNeurIPS 2024 · 305 citations
- Unveiling Factual Recall Behaviors of Large Language Models through Knowledge NeuronsYifei Wang, Yuheng Chen, Wanting Wen, Yu Sheng et al.EMNLP 2024 · 3 citations
- DeCoT: Debiasing Chain-of-Thought for Knowledge-Intensive Tasks in Large Language Models via Causal InterventionJunda Wu, Tong Yu, Xiang Chen, Haoliang Wang et al.ACL 2024
- Complexity-Based Prompting for Multi-step ReasoningYao Fu, Hao Peng, Ashish Sabharwal, Peter Clark et al.ICLR 2023 · 73 citations
- The Potential of CoT for Reasoning: A Closer Look at Trace DynamicsGregor Bachmann, Yichen Jiang, Seyed-Mohsen Moosavi-Dezfooli, Moin NabiICLR 2026 · 5 citations
