PDTrim: Targeted Pruning for Prefill-Decode Disaggregation in Inference
Hao Zhang, Mengsi Lyu, Zhuo Chen, Yulong Ao, Yonghua Lin
Abstract
Large Language Models (LLMs) demonstrate exceptional capabilities across various tasks, but their deployment is constrained by high computational and memory costs. Model pruning provides an effective means to alleviate these demands. However, existing methods often ignore the characteristics of prefill-decode (PD) disaggregation in practice. In this paper, we propose a pruning method that is highly integrated with PD disaggregation, enabling more precise pruning of blocks. Our approach constructs pruning and distillation sets to perform iterative block removal, obtaining better pruning solutions. Moreover, we analyze the pruning sensitivity of the prefill and decode stages and identify removable blocks specific to each stage, making it well suited for PD disaggregation deployment. Extensive experiments demonstrate our approach consistently achieves strong performance in both PD disaggregation and PD unified (non-PD disaggregation) settings, and can also be extended to other non-block pruning methods. Under the same settings, our method achieves improved performance and faster inference.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4cf5ac44-3ff0-452f-a10d-4f1485f9cea6Cited by top-tier papers2
- ShieldedCode: Learning Robust Representations for Virtual Machine Protected CodeMingqiao Mo, Yunlong Tan, Hao Zhang, Heng Zhang et al.ICLR 2026 · 10 citations
- AEA: Adaptive Expert Allocation Improves Sentence Embeddings from Mixture-of-Experts LLMShufan Yang, Zifeng Cheng, Zhiwei Jiang, Qingfeng Qi et al.ACL 2026
Builds on13
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language ModelsGuangxuan Xiao, Ji Lin, Mickaël Seznec, Hao Wu et al.ICML 2023 · 1,493 citations
- DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model ServingYinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu et al.OSDI 2024 · 646 citations
- Post-Training Quantization for Vision TransformerZhenhua Liu, Yunhe Wang, Kai Han, Wei Zhang et al.NeurIPS 2021 · 528 citations
Related papers
- Efficient Multi-round LLM Inference over Disaggregated ServingWenhao He, Youhe Jiang, Penghao Zhao, Quanqing Xu et al.ICML 2026 · 7 citations
- Efficiently Serving Large Multimodal Models Using EPD DisaggregationGursimran Singh, Xinglu Wang, Yifan Hu, Timothy Tin Long Yu et al.ICML 2025
- Layer as Puzzle Pieces: Compressing Large Language Models through Layer ConcatenationFei Wang, Li Shen, Liang Ding, Chao Xue et al.NeurIPS 2025 · 7 citations
- Let LLM Tell What to Prune and How Much to PruneMingzhe Yang, Sihao Lin, Changlin Li, Xiaojun ChangICML 2025
- Gradient-based Intra-attention Pruning on Pre-trained Language ModelsZiqing Yang, Yiming Cui, Xin Yao, Shijin WangACL 2023 · 2 citations
