An Adaptive Language-Agnostic Pruning Method for Greener Language Models for Code
Mootez Saad, José Antonio Hernández López, Boqi Chen, Dániel Varró, Tushar Sharma
Abstract
Language models of code have demonstrated remarkable performance across various software engineering and source code analysis tasks. However, their demanding computational resource requirements and consequential environmental footprint remain as significant challenges. This work introduces Alpine, an adaptive programming language-agnostic pruning technique designed to substantially reduce the computational overhead of these models. The proposed method offers a pluggable layer that can be integrated with all Transformer-based models. With Alpine, input sequences undergo adaptive compression throughout the pipeline, reaching a size that is up to ×3 less their initial size, resulting in significantly reduced computational load. Our experiments on two software engineering tasks, defect prediction and code clone detection across three language models CodeBert, GraphCodeBert and UniXCoder show that Alpine achieves up to a 50% reduction in FLOPs, a 58.1% decrease in memory footprint, and a 28.1% improvement in throughput on average. This led to a reduction in CO 2 emissions by up to 44.85%. Importantly, it achieves a reduction in computation resources while maintaining up to 98.1% of the original predictive performance. These findings highlight the potential of Alpine in making language models of code more resource-efficient and accessible while preserving their performance, contributing to the overall sustainability of their adoption in software development. Also, it sheds light on redundant and noisy information in source code analysis corpora, as shown by the substantial sequence compression achieved by Alpine.
CCS Concepts: • Software and its engineering → Software notations and tools; • Computing methodologies → Machine learning algorithms;
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e8aaebc1-480a-4432-84dd-cb99b5133e24Cited by top-tier papers2
- Reducing Cost of LLM Agents with Trajectory ReductionYuan-An Xiao, Pengfei Gao, Chao Peng, Yingfei XiongFSE 2026 · 1 citation
- Carbon-Taxed Transformers: A Green Compression Pipeline for Overgrown Language ModelsAjmain Inqiad Alam, Palash Ranjan Roy, Chanchal K. Roy, Banani Roy et al.FSE 2026
Builds on20
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- GraphCodeBERT: Pre-training Code Representations with Data FlowDaya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng et al.ICLR 2021 · 1,644 citations
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 1,224 citations
- ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale TransformersZhewei Yao, Reza Yazdani Aminabadi, Minjia Zhang, Xiaoxia Wu et al.NeurIPS 2022 · 816 citations
- DS-1000: A Natural and Reliable Benchmark for Data Science Code GenerationYuhang Lai, Chengxi Li, Yiming Wang, Tianyi Zhang et al.ICML 2023 · 504 citations
Related papers
- Compressing Pre-trained Models of Code into 3 MBJieke Shi, Zhou Yang, Bowen Xu, Hong Jin Kang et al.ASE 2022 · 42 citations
- An Empirical Study of Parameter-Efficient Fine-Tuning Methods for Pre-Trained Code ModelsJiaxing Liu, Chaofeng Sha, Xin PengASE 2023 · 24 citations
- Diet code is healthy: simplifying programs for pre-trained models of codeZhaowei Zhang, Hongyu Zhang, Beijun Shen, Xiaodong GuFSE 2022 · 39 citations
- COPAL: Continual Pruning in Large Language Generative ModelsSrikanth Malla, Joon Hee Choi, Chiho ChoiICML 2024 · 6 citations
- LLM-Pruner: On the Structural Pruning of Large Language ModelsXinyin Ma, Gongfan Fang, Xinchao WangNeurIPS 2023 · 994 citations
