Looking Beyond the Top-1: Transformers Determine Top Tokens in Order
Daria Lioubashevski, Tomer Schlank, Gabriel Stanovsky, Ariel Goldstein
Abstract
Uncovering the inner mechanisms of Transformer models offers insights into how they process and represent information. In this work, we analyze the computation performed by Transformers in the layers after the top-1 prediction remains fixed, known as the "saturation event". We expand this concept to top-k tokens, demonstrating that similar saturation events occur across language, vision, and speech models. We find that these events occur in order of the corresponding tokens' ranking, i.e., the model first decides on the top ranking token, then the second highest ranking token, and so on. This phenomenon seems intrinsic to the Transformer architecture, occurring across different variants, and even in untrained Transformers. We propose that these events reflect task transitions, where determining each token corresponds to a discrete task. We show that it is possible to predict the current task from hidden layer embedding, and demonstrate that we can cause the model to switch to the next task via intervention. Leveraging our findings, we introduce a tokenlevel early-exit strategy, surpassing existing methods in balancing performance and efficiency and show how to exploit saturation events for better language modeling.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical FindingsQiong Wu, Wenhao Lin, Yiyi Zhou, Weihao Ye et al.NeurIPS 2025 · 16 citations
- LUMINA: Detecting Hallucinations in RAG System with Context–Knowledge SignalsSamuel Yeh, Sharon Li, Tanwi MallickICLR 2026 · 11 citations
- Rethinking Layer Relevance in Large Language Models Beyond Cosine SimilarityCristian Hinostroza, Rodrigo Toro Icarte, Christ Devia, Andres Carvallo et al.ICLR 2026 · 4 citations
- Inverse Depth Scaling From Most Layers Being SimilarYizhou Liu, Sara Kangaslahti, Ziming Liu, Jeff GoreICML 2026 · 4 citations
Builds on18
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- Confident Adaptive Language ModelingTal Schuster, Adam Fisch, Jai Gupta, Mostafa Dehghani et al.NeurIPS 2022 · 394 citations
Related papers
- Tracing Representation Progression: Analyzing and Enhancing Layer-Wise SimilarityJiachen Jiang, Jinxin Zhou, Zhihui ZhuICLR 2025
- Abrupt Learning in Transformers: A Case Study on Matrix CompletionPulkit Gopalani, Ekdeep Singh Lubana, Wei HuNeurIPS 2024 · 12 citations
- Leveraging Relaxed Equilibrium by Lazy Transition for Sequence ModelingXi Ai, Bin FangACL 2022
- Eureka-Moments in Transformers: Multi-Step Tasks Reveal Softmax Induced Optimization ProblemsDavid T. Hoffmann, Simon Schrodi, Jelena Bratulic, Nadine Behrmann et al.ICML 2024 · 11 citations
- Attention Sinks and Compression Valleys in LLMs are Two Sides of the Same CoinEnrique Queipo-de-Llano, Alvaro Arroyo, Federico Barbero, Xiaowen Dong et al.ICLR 2026 · 56 citations
