FFN-SkipLLM: A Hidden Gem for Autoregressive Decoding with Adaptive Feed Forward Skipping
Ajay Jaiswal, Bodun Hu, Lu Yin, Yeonju Ro, Tianlong Chen, Shiwei Liu, Aditya Akella
Abstract
Autoregressive Large Language Models (e.g., LLaMa, GPTs) are omnipresent achieving remarkable success in language understanding and generation. However, such impressive capability typically comes with a substantial model size, which presents significant challenges for autoregressive token-by-token generation. To mitigate computation overload incurred during generation, several early-exit and layer-dropping strategies have been proposed. Despite some promising success due to the redundancy across LLMs layers on metrics like Rough-L/BLUE, our careful knowledgeintensive evaluation unveils issues such as generation collapse, hallucination, and noticeable performance drop even at the trivial exit ratio of ∼ 10-15% of layers. We attribute these errors primarily to ineffective handling of the KV cache through state copying during early exit. In this work, we observe the saturation of computationally expensive feed-forward blocks of LLM layers and propose FFN-SkipLLM, which is a novel fine-grained skip strategy for autoregressive LLMs. FFN-SkipLLM leverages an input-adaptive feed-forward skipping approach that can skip ∼ 25-30% of FFN blocks of LLMs with marginal change in performance on knowledge-intensive generation tasks without any requirement to handle the KV cache. Our extensive experiments and ablation studies across benchmarks like MT-Bench, Factoid-QA, and variable-length text summarization illustrate how our simple and easy-touse method can facilitate faster autoregressive decoding. PROMPT >> Please provide answer to the following. Question: Who is the prime minister of India? SkipDecode Assistant 25% Hello! India currently does not have a prime minister since India abolished its cabinet posts including Prime Minister Narendra Mod Mod Prime Minister Mod Mod resigned as Prime Minister of India effective immediately after his party losts seats in parliamentary elections held earlier this month. India now transitioning into transition mode transition mode transition mode transition mode transition mode .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- AdaSkip: Adaptive Sublayer Skipping for Accelerating Long-Context LLM InferenceZhuomin He, Yizhen Yao, Pengfei Zuo, Bin Gao et al.AAAI 2025 · 13 citations
- MemoryLLM: Plug-n-Play Interpretable Feed-Forward Memory for TransformersAjay Jaiswal, Lauren Hannah, Han-Byul Kim, Duc Hoang et al.ICML 2026 · 3 citations
- SwiftKV: Fast Prefill-Optimized Inference with Knowledge-Preserving Model TransformationAurick Qiao, Zhewei Yao, Samyam Rajbhandari, Yuxiong HeEMNLP 2025 · 1 citation
- NoVo: Norm Voting off Hallucinations with Attention Heads in Large Language ModelsZheng Yi Ho, Siyuan Liang, Sen Zhang, Yibing Zhan et al.ICLR 2025
- StitchLLM: Serving LLMs, One Block at a TimeBodun Hu, Shuozhe Li, Saurabh Agarwal, Myungjin Lee et al.ACL 2025
Builds on22
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
- SparseGPT: Massive Language Models Can be Accurately Pruned in One-ShotElias Frantar, Dan AlistarhICML 2023 · 1,240 citations
- VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric TasksWenhai Wang, Zhe Chen, Xiaokang Chen, Jiannan Wu et al.NeurIPS 2023 · 725 citations
- Chameleon: Plug-and-Play Compositional Reasoning with Large Language ModelsPan Lu, Baolin Peng, Hao Cheng, Michel Galley et al.NeurIPS 2023 · 515 citations
- Is ChatGPT a General-Purpose Natural Language Processing Task Solver?Chengwei Qin, Aston Zhang, Zhuosheng Zhang, Jiaao Chen et al.EMNLP 2023 · 449 citations
Related papers
- MoDES: Accelerating Mixture-of-Experts Multimodal Large Language Models via Dynamic Expert SkippingYushi Huang, Zining Wang, Zhihang Yuan, Yifu Ding et al.CVPR 2026 · 15 citations
- VideoLLM-MoD: Efficient Video-Language Streaming with Mixture-of-Depths Vision ComputationShiwei Wu, Joya Chen, Kevin Qinghong Lin, Qimeng Wang et al.NeurIPS 2024 · 78 citations
- Chain-of-Thought Compression Should Not Be Blind: V-Skip for Efficient Multimodal Reasoning via Dual-Path AnchoringDongxu Zhang, Yiding Sun, Cheng Tan, Wenbiao Yan et al.ACL 2026 · 18 citations
- Skip-Vision: Efficient and Scalable Acceleration of Vision-Language Models Via Adaptive Token SkippingWeili Zeng, Ziyuan Huang, Kaixiang Ji, Yichao YanICCV 2025
- LayerSkip: Enabling Early Exit Inference and Self-Speculative DecodingMostafa Elhoushi, Akshat Shrivastava, Diana Liskovich, Basil Hosmer et al.ACL 2024 · 22 citations
