D-LLM: A Token Adaptive Computing Resource Allocation Strategy for Large Language Models
Yikun Jiang, Huanyu Wang, Lei Xie, Hanbin Zhao, Zhang Chao, Hui Qian, John C. S. Lui
Abstract
Large language models have shown an impressive societal impact owing to their excellent understanding and logical reasoning skills. However, such strong ability relies on a huge amount of computing resources, which makes it difficult to deploy LLMs on computing resource-constrained platforms. Currently, LLMs process each token equivalently, but we argue that not every word is equally important. Some words should not be allocated excessive computing resources, particularly for dispensable terms in simple questions. In this paper, we propose a novel dynamic inference paradigm for LLMs, namely D-LLMs, which adaptively allocate computing resources in token processing. We design a dynamic decision module for each transformer layer that decides whether a network unit should be executed or skipped. Moreover, we tackle the issue of adapting D-LLMs to real-world applications, specifically concerning the missing KV-cache when layers are skipped. To overcome this, we propose a simple yet effective eviction policy to exclude the skipped layers from subsequent attention calculations. The eviction policy not only enables D-LLMs to be compatible with prevalent applications but also reduces considerable storage resources. Experimentally, D-LLMs show superior performance, in terms of computational cost and KV storage utilization. It can reduce up to 45% computational cost and KV storage on Q&A, summarization, and math solving tasks, 50% on commonsense reasoning tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers13
- Retrieval-of-Thought: Efficient Reasoning via Reusing ThoughtsAmmar Ahmed, Azal Ahmad Khan, Ayaan Ahmad, Sheng Di et al.ICLR 2026 · 13 citations
- SpecEE: Accelerating Large Language Model Inference with Speculative Early ExitingJiaming Xu, Jiayi Pan, Yongkang Zhou, Siming Chen et al.ISCA 2025 · 9 citations
- EfficientNav: Towards On-Device Object-Goal Navigation with Navigation Map Caching and RetrievalZebin Yang, Sunjian Zheng, Tong Xie, Tianshi Xu et al.NeurIPS 2025 · 7 citations
- Learning When to Attend: Conditional Memory Access for Long-Context LLMsSakshi Choudhary, Aditya Chattopadhyay, Luca Zancato, Elvis Nunez et al.ICML 2026 · 2 citations
- Equilibrium Language ModelsYikun Jiang, Huanyu Wang, Tianhong Ding, Wenhu Zhang et al.ICLR 2026
Builds on32
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
Related papers
- Sparse-dLLM: Accelerating Diffusion LLMs with Dynamic Cache EvictionYuerong Song, Xiaoran Liu, Ruixiao Li, Zhigeng Liu et al.AAAI 2026 · 43 citations
- Dynamic Thinking-Token Selection for Efficient Reasoning in Large Reasoning Modelszhenyuan guo, Tong Chen, Wenlong Meng, Chen GONG et al.ICML 2026 · 1 citation
- LazyEviction: Lagged KV Eviction with Attention Pattern Observation for Efficient Long ReasoningHaoyue Zhang, Hualei Zhang, Xiaosong Ma, Jie Zhang et al.ACL 2026 · 7 citations
- Skip a Layer or Loop It? Learning Program-of-Layers in LLMsZiyue Li, Yang Li, Tianyi ZhouICML 2026 · 4 citations
- Inference-Time Hyper-Scaling with KV Cache CompressionAdrian Lancucki, Konrad Staniszewski, Piotr Nawrot, Edoardo Maria PontiNeurIPS 2025 · 36 citations
