What Layers When: Learning to Skip Compute in LLMs with Residual Gates
Filipe Laitenberger, Dawid Jan Kopiczko, Cees G. M. Snoek, Yuki M. Asano
摘要
We introduce GateSkip, a simple residual-stream gating mechanism that enables token-wise layer skipping in decoder-only LMs. Each Attention/MLP branch is equipped with a sigmoid-linear gate that condenses the branch's output before it re-enters the residual stream. During inference we rank tokens by the gate values and skip low-importance ones using a per-layer budget. While early-exit or router-based Mixture-of-Depths models are known to be unstable and need extensive retraining, our smooth, differentiable gates fine-tune stably on top of pretrained models. On long-form reasoning, we save up to 15% compute while retaining over 90% of baseline accuracy. For increasingly larger models, this tradeoff improves drastically. On instruction-tuned models we see accuracy gains at full compute and match baseline quality near 50% savings. The learned gates give insight into transformer information flow (e.g., BOS tokens act as anchors), and the method combines easily with quantization, pruning, and self-speculative decoding.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper18
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 被引用 3,037 次
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao 等AAAI 2020 · 被引用 2,916 次
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han 等ICLR 2024 · 被引用 1,714 次
- Vision Transformers Need RegistersTimothée Darcet, Maxime Oquab, Julien Mairal, Piotr BojanowskiICLR 2024 · 被引用 769 次
相关 Paper
- TokenSkip: Controllable Chain-of-Thought Compression in LLMsHeming Xia, Chak Tou Leong, Wenjie Wang, Yongqi Li 等EMNLP 2025 · 被引用 6 次
- MoDES: Accelerating Mixture-of-Experts Multimodal Large Language Models via Dynamic Expert SkippingYushi Huang, Zining Wang, Zhihang Yuan, Yifu Ding 等CVPR 2026 · 被引用 15 次
- LayerSkip: Enabling Early Exit Inference and Self-Speculative DecodingMostafa Elhoushi, Akshat Shrivastava, Diana Liskovich, Basil Hosmer 等ACL 2024 · 被引用 22 次
- SkipGPT: Each Token is One of a KindAnhao Zhao, Fanghua Ye, Yingqi Fan, Junlong Tong 等ICML 2025
- Router-Tuning: A Simple and Effective Approach for Dynamic DepthShwai He, Tao Ge, Guoheng Sun, Bowei Tian 等EMNLP 2025
