TrimLLM: Progressive Layer Dropping for Domain-Specific LLMs
Lanxiang Hu, Tajana Rosing, Hao Zhang
摘要
Specializing large language models (LLMs) for local deployment in domain-specific use cases is necessary for strong performance while meeting latency and privacy constraints. However, conventional task-specific adaptation approaches do not show simultaneous memory saving and inference speedup at deployment time. Practical compression techniques like quantization and pruning require dedicated hardware or kernel support to achieve measured inference speedup. We develop TRIM-LLM based on the layer-wise specialization phenomenon we empirically observed and verified on contemporary LLMs. TRIMLLM reduces the depth of LLMs via progressive layer dropping. We show it retains LLMs' capacity in specific domains and achieves inference speedup irrespective of hardware and deep learning frameworks. We evaluated TRIMLLM on LLMs of various sizes for inference; models adapted on medical, legal, and financial datasets all demonstrate 2.1 -5.7× inference speedup on consumer GPUs and up to 3.1× speedup on A100 when compared to state-ofthe-art model compression algorithms, with no loss in accuracy at 50∼60% model compression ratio.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- EfficientVLA: Training-Free Acceleration and Compression for Vision-Language-Action ModelsYantai Yang, Yuhao Wang, Zichen Wen, Luo Zhongwei 等NeurIPS 2025 · 被引用 94 次
- MARD: Module-Aware Reasoning Distillation for Language Models with Adaptive SupervisionWenqi Yang, Jianjun Li, Zhibo Zhang, Mingqian Ding 等ACL 2026
它引用的顶会 Paper15
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao 等AAAI 2020 · 被引用 2,916 次
- SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language ModelsGuangxuan Xiao, Ji Lin, Mickaël Seznec, Hao Wu 等ICML 2023 · 被引用 1,493 次
- LLM-Pruner: On the Structural Pruning of Large Language ModelsXinyin Ma, Gongfan Fang, Xinchao WangNeurIPS 2023 · 被引用 994 次
- A Simple and Effective Pruning Approach for Large Language ModelsMingjie Sun, Zhuang Liu, Anna Bair, J. Zico KolterICLR 2024 · 被引用 794 次
相关 Paper
- EDGE-LLM: Enabling Efficient Large Language Model Adaptation on Edge Devices via Unified Compression and Adaptive Layer VotingZhongzhi Yu, Zheng Wang, Yuhan Li, Ruijie Gao 等DAC 2024 · 被引用 57 次
- ZipLM: Inference-Aware Structured Pruning of Language ModelsEldar Kurtic, Elias Frantar, Dan AlistarhNeurIPS 2023 · 被引用 69 次
- Let LLM Tell What to Prune and How Much to PruneMingzhe Yang, Sihao Lin, Changlin Li, Xiaojun ChangICML 2025
- MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language ModelsElias Frantar, Roberto L. Castro, Jiale Chen, Torsten Hoefler 等PPoPP 2025 · 被引用 24 次
- Radio: Rate-Distortion Optimization for Large Language Model CompressionSean I. YoungICML 2025
