APT: Adaptive Pruning and Tuning Pretrained Language Models for Efficient Training and Inference
Bowen Zhao, Hannaneh Hajishirzi, Qingqing Cao
Abstract
Fine-tuning and inference with large Language Models (LM) are generally known to be expensive. Parameter-efficient fine-tuning over pretrained LMs reduces training memory by updating a small number of LM parameters but does not improve inference efficiency. Structured pruning improves LM inference efficiency by removing consistent parameter blocks, yet often increases training memory and time. To improve both training and inference efficiency, we introduce APT that adaptively prunes and tunes parameters for the LMs. At the early stage of finetuning, APT dynamically adds salient tuning parameters for fast and accurate convergence while discarding unimportant parameters for efficiency. Compared to baselines, our experiments show that APT maintains up to 98% task performance when pruning 60% of the parameters in RoBERTa and T5 models. APT also preserves 86.4% of LLaMA models' performance with 70% parameters remaining. Furthermore, APT speeds up LMs' fine-tuning by up to 8× and reduces large LMs' memory training footprint by up to 70%. Our code and models are publicly available at https://github.com/ROIM1998/APT .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b4b7e595-dc91-4352-b1ec-f5ba177161edCited by top-tier papers12
- Unveiling and Harnessing Hidden Attention Sinks: Enhancing Large Language Models without Training through Attention CalibrationZhongzhi Yu, Zheng Wang, Yonggan Fu, Huihong Shi et al.ICML 2024 · 63 citations
- LoRAFusion: Efficient LoRA Fine-Tuning for LLMsZhanda Zhu, Qidong Su, Yaoyao Ding, Kevin Song et al.EuroSys 2026 · 2 citations
- Lua-LLM: Learning Unstructured-Sparsity Allocation for Large Language ModelsMingge Lu, Jingwei Sun, Junqing Lin, Zechun Zhou et al.NeurIPS 2025 · 1 citation
- MaskPro: Linear-Space Probabilistic Learning for Strict (N:M)-Sparsity on LLMsYan Sun, Qixin Zhang, Zhiyuan Yu, Xikun Zhang et al.ICLR 2026 · 1 citation
- Computation and Memory-Efficient Model Compression with Gradient ReweightingZhiwei Li, Yuesen Liao, Binrui Wu, Yuquan Zhou et al.NeurIPS 2025
Builds on23
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 3,228 citations
- Towards a Unified View of Parameter-Efficient Transfer LearningJunxian He, Chunting Zhou, Xuezhe Ma, Taylor Berg-Kirkpatrick et al.ICLR 2022 · 1,182 citations
- Cross-Task Generalization via Natural Language Crowdsourcing InstructionsSwaroop Mishra, Daniel Khashabi, Chitta Baral, Hannaneh HajishirziACL 2022 · 887 citations
Related papers
- PAT: Pruning-Aware Tuning for Large Language ModelsYijiang Liu, Huanrui Yang, Youxin Chen, Rongyu Zhang et al.AAAI 2025 · 1 citation
- Sketch to Adapt: Fine-Tunable Sketches for Efficient LLM AdaptationTianyi Zhang, Junda Su, Aditya Desai, Oscar Wu et al.ICML 2025
- TARE: Lightweight Token-Aware Representation Editing for Fine-tuning Transformer-like ModelsYulong Wang, Siyu ZhaoACL 2026
- DLP: Dynamic Layerwise Pruning in Large Language ModelsYuli Chen, Bo Cheng, Jiale Han, Yingying Zhang et al.ICML 2025
- From Bottom to Top: Extending the Potential of Parameter Efficient Fine-TuningJihao Gu, Zelin Wang, Yibo Zhang, Ziji Zhang et al.EMNLP 2024 · 3 citations
