Leap-of-Thought: Accelerating Transformers via Dynamic Token Routing
Yeachan Kim, Junho Kim, Jun-Hyung Park, Mingyu Lee, SangKeun Lee
Abstract
Computational inefficiency in transformers has been a long-standing challenge, hindering the deployment in resource-constrained or realtime applications. One promising approach to mitigate this limitation is to progressively remove less significant tokens, given that the sequence length strongly contributes to the inefficiency. However, this approach entails a potential risk of losing crucial information due to the irrevocable nature of token removal. In this paper, we introduce Leap-of-Thought (LoT), a novel token reduction approach that dynamically routes tokens within layers. Unlike previous work that irrevocably discards tokens, LoT enables tokens to 'leap' across layers. This ensures that all tokens remain accessible in subsequent layers while reducing the number of tokens processed within layers. We achieve this by pairing the transformer with dynamic token routers, which learn to selectively process tokens essential for the task. Evaluation results clearly show that LoT achieves substantial improvement on computational efficiency. Specifically, LoT attains up to 25× faster inference time without a significant loss in accuracy 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on16
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- Reducing Transformer Depth on Demand with Structured DropoutAngela Fan, Edouard Grave, Armand JoulinICLR 2020 · 695 citations
Related papers
- Learned Token Pruning for TransformersSehoon Kim, Sheng Shen, David Thorsley, Amir Gholami et al.KDD 2022 · 97 citations
- Bring Future Vision: Dynamic Computation Allocation Guided by Lightweight Feature ForecasterChao Han, Yijuan Liang, Zihao Xuan, Daokuan Wu et al.ICML 2026
- Dynamic Context Pruning for Efficient and Interpretable Autoregressive TransformersSotiris Anagnostidis, Dario Pavllo, Luca Biggio, Lorenzo Noci et al.NeurIPS 2023 · 95 citations
- Transkimmer: Transformer Learns to Layer-wise SkimYue Guan, Zhengyi Li, Jingwen Leng, Zhouhan Lin et al.ACL 2022
- Token Dropping for Efficient BERT PretrainingLe Hou, Richard Yuanzhe Pang, Tianyi Zhou, Yuexin Wu et al.ACL 2022
