CoTFormer: A Chain of Thought Driven Architecture with Budget-Adaptive Computation Cost at Inference
Amirkeivan Mohtashami, Matteo Pagliardini, Martin Jaggi
摘要
Scaling language models to larger and deeper sizes has led to significant boosts in performance. Even though the size of these models limits their application in compute-constrained environments, the race to continually develop ever larger and deeper foundational models is underway. At the same time-regardless of the model size-task-specific techniques continue to play a pivotal role in achieving optimal downstream performance. One of these techniques, called Chain-of-Thought (CoT), is particularly interesting since, as we point out in this work, it resembles employing a deeper transformer through re-applying the model multiple times. However, a key subtlety in computing the attention of past tokens differentiates CoT from simply applying the model several times. Based on this insight, we propose CoTFormer, a novel architecture which closely mimics CoT at the token level, allowing us to obtain significantly improved accuracies close to much larger models. While applying CoT introduces additional computation costs, we compensate for it by leveraging CoTFormer's special compatibility with token-wise variable depth. Through a compute adaptive model-which automatically allocates the compute to tokens that need it most-we show that it is possible to reduce the computation cost significantly without any reduction in accuracy, and with further compute cost reductions possible while maintaining a competitive accuracy. * Equal contribution. order is alphabetical.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- LaDiR: Latent Diffusion Enhances LLMs for Text ReasoningHaoqiang Kang, Yizhe Zhang, Nikki Lijing Kuang, Nicklas Majamaki 等ICLR 2026 · 被引用 25 次
- Coevolutionary Continuous Discrete Diffusion: Make Your Diffusion Language Model a Latent ReasonerCai Zhou, Chenxiao Yang, Yi Hu, Chenyu Wang 等ICML 2026 · 被引用 21 次
- MeSH: Memory-as-State-Highways for Recursive TransformersChengting Yu, Xiaobo Shu, Yadao Wang, Yizhen Zhang 等ICLR 2026 · 被引用 10 次
- Stabilizing Recurrent Dynamics for Test-Time Scalable Latent Reasoning in Looped Language ModelsXiao-Wen Yang, Zi-Yu Han, Xi-Hua Zhang, Wen-Da Wei 等ICML 2026 · 被引用 6 次
- Efficient Parallel Samplers for Recurrent-Depth ModelsJonas Geiping, Xinyu Yang, Guinan SuICML 2026 · 被引用 5 次
它引用的顶会 Paper13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- On Layer Normalization in the Transformer ArchitectureRuibin Xiong, Yunchang Yang, Di He, Kai Zheng 等ICML 2020 · 被引用 1,388 次
相关 Paper
- Chain of Thought Empowers Transformers to Solve Inherently Serial ProblemsZhiyuan Liu, Hong Liu, Denny Zhou, Tengyu MaICLR 2024 · 被引用 259 次
- LoopFormer: Elastic-Depth Looped Transformers for Latent Reasoning via Shortcut ModulationAhmadreza Jeddi, Marco Ciccone, Babak TaatiICLR 2026 · 被引用 54 次
- Iteration Head: A Mechanistic Study of Chain-of-ThoughtVivien Cabannes, Charles Arnal, Wassim Bouaziz, Xingyu Yang 等NeurIPS 2024 · 被引用 44 次
- Towards Revealing the Mystery behind Chain of Thought: A Theoretical PerspectiveGuhao Feng, Bohang Zhang, Yuntian Gu, Haotian Ye 等NeurIPS 2023 · 被引用 470 次
- From Sparse Dependence to Sparse Attention: Unveiling How Chain-of-Thought Enhances Transformer Sample EfficiencyKaiyue Wen, Huaqing Zhang, Hongzhou Lin, Jingzhao ZhangICLR 2025
