Diet code is healthy: simplifying programs for pre-trained models of code
Zhaowei Zhang, Hongyu Zhang, Beijun Shen, Xiaodong Gu
Abstract
Pre-trained code representation models such as CodeBERT have demonstrated superior performance in a variety of software engineering tasks, yet they are often heavy in complexity, quadratically with the length of the input sequence. Our empirical analysis of CodeBERT's attention reveals that CodeBERT pays more attention to certain types of tokens and statements such as keywords and data-relevant statements. Based on these findings, we propose DietCode, which aims at lightweight leverage of large pre-trained models for source code. DietCode simplifies the input program of CodeBERT with three strategies, namely, word dropout, frequency filtering, and an attention-based strategy that selects statements and tokens that receive the most attention weights during pre-training. Hence, it gives a substantial reduction in the computational cost without hampering the model performance. Experimental results on two downstream tasks show that DietCode provides comparable results to CodeBERT with 40% less computational cost in fine-tuning and testing.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3fc0507f-ce58-40a4-9da0-881bb3cee099Cited by top-tier papers20
- NatGen: generative pre-training by "naturalizing" source codeSaikat Chakraborty, Toufique Ahmed, Yangruibo Ding, Premkumar T. Devanbu et al.FSE 2022 · 101 citations
- An Empirical Comparison of Pre-Trained Models of Source CodeChangan Niu, Chuanyi Li, Vincent Ng, Dongxiao Chen et al.ICSE 2023 · 71 citations
- CCTEST: Testing and Repairing Code Completion SystemsZongjie Li, Chaozheng Wang, Zhibo Liu, Haoxuan Wang et al.ICSE 2023 · 49 citations
- Towards Understanding the Characteristics of Code Generation Errors Made by Large Language ModelsZhijie Wang, Zijie Zhou, Da Song, Yuheng Huang et al.ICSE 2025 · 12 citations
- AI Coders Are among Us: Rethinking Programming Language Grammar towards Efficient Code GenerationZhensu Sun, Xiaoning Du, Zhou Yang, Li Li et al.ISSTA 2024 · 8 citations
Builds on11
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 1,224 citations
- Q-BERT: Hessian Based Ultra Low Precision Quantization of BERTSheng Shen, Zhen Dong, Jiayu Ye, Linjian Ma et al.AAAI 2020 · 656 citations
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 541 citations
- NatGen: generative pre-training by "naturalizing" source codeSaikat Chakraborty, Toufique Ahmed, Yangruibo Ding, Premkumar T. Devanbu et al.FSE 2022 · 101 citations
Related papers
- Natural Is the Best: Model-Agnostic Code Simplification for Pre-trained Large Language ModelsYan Wang, Xiaoning Li, Tien N. Nguyen, Shaohua Wang et al.FSE 2024 · 6 citations
- Towards Efficient Fine-Tuning of Pre-trained Code Models: An Experimental Study and BeyondEnsheng Shi, Yanlin Wang, Hongyu Zhang, Lun Du et al.ISSTA 2023 · 37 citations
- Compressing Pre-trained Models of Code into 3 MBJieke Shi, Zhou Yang, Bowen Xu, Hong Jin Kang et al.ASE 2022 · 42 citations
- GraphCodeBERT: Pre-training Code Representations with Data FlowDaya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng et al.ICLR 2021 · 1,644 citations
- LEANCODE: Understanding Models Better for Code Simplification of Pre-trained Large Language ModelsYan Wang, Ling Ding, Tien N. Nguyen, Shaohua Wang et al.ACL 2025 · 1 citation
