Prompt-based Depth Pruning of Large Language Models
Juyun Wee, Minjae Park, Jaeho Lee
Abstract
Depth pruning aims to reduce the inference cost of a large language model without any hardwarespecific complications, by simply removing several less important transformer blocks. However, our empirical findings suggest that the importance of a transformer block may be highly taskdependent-a block that is crucial for a task can be removed without degrading the accuracy on another task. Based on this observation, we develop a dynamic depth pruning algorithm, coined PuD-Ding (Prompt-routed Dynamic Depth Pruning), which determines which blocks to omit from the model based on the input prompt. PuDDing operates by training a lightweight router to predict the best omission set among a set of options, where this option set has also been constructed in a data-driven manner. Empirical results on commonsense reasoning benchmarks demonstrate that PuDDing effectively accelerates the inference language models, and achieves better on-task performance than static depth pruning baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- SubspacePath Pruner: Inference-time Pruning via Probe-based Representation–Parameter CouplingZhiren Gong, Yikun Hou, Fan Wu, CHE WANG et al.ICML 2026
- OCP: Outlier-Centric Probing for Dynamic Structured Pruning of LLMsYang Ji, Ying SunACL 2026
Builds on10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- Deja Vu: Contextual Sparsity for Efficient LLMs at Inference TimeZichang Liu, Jue Wang, Tri Dao, Tianyi Zhou et al.ICML 2023 · 318 citations
Related papers
- IG-Pruning: Input-Guided Block Pruning for Large Language ModelsKangyu Qiao, Shaolei Zhang, Yang FengEMNLP 2025 · 1 citation
- Let LLM Tell What to Prune and How Much to PruneMingzhe Yang, Sihao Lin, Changlin Li, Xiaojun ChangICML 2025
- Dr.LLM: Dynamic Layer Routing in LLMsAhmed Heakl, Martin Gubri, Salman Khan, Sangdoo Yun et al.ICLR 2026 · 11 citations
- Probe Pruning: Accelerating LLMs through Dynamic Pruning via Model-ProbingQi Le, Enmao Diao, Ziyan Wang, Xinran Wang et al.ICLR 2025
- Stop Looking for "Important Tokens" in Multimodal Language Models: Duplication Matters MoreZichen Wen, Yifeng Gao, Shaobo Wang, Junyuan Zhang et al.EMNLP 2025 · 4 citations
