Diet code is healthy: simplifying programs for pre-trained models of code
Zhaowei Zhang, Hongyu Zhang, Beijun Shen, Xiaodong Gu
摘要
Pre-trained code representation models such as CodeBERT have demonstrated superior performance in a variety of software engineering tasks, yet they are often heavy in complexity, quadratically with the length of the input sequence. Our empirical analysis of CodeBERT's attention reveals that CodeBERT pays more attention to certain types of tokens and statements such as keywords and data-relevant statements. Based on these findings, we propose DietCode, which aims at lightweight leverage of large pre-trained models for source code. DietCode simplifies the input program of CodeBERT with three strategies, namely, word dropout, frequency filtering, and an attention-based strategy that selects statements and tokens that receive the most attention weights during pre-training. Hence, it gives a substantial reduction in the computational cost without hampering the model performance. Experimental results on two downstream tasks show that DietCode provides comparable results to CodeBERT with 40% less computational cost in fine-tuning and testing.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper20
- NatGen: generative pre-training by "naturalizing" source codeSaikat Chakraborty, Toufique Ahmed, Yangruibo Ding, Premkumar T. Devanbu 等FSE 2022 · 被引用 101 次
- An Empirical Comparison of Pre-Trained Models of Source CodeChangan Niu, Chuanyi Li, Vincent Ng, Dongxiao Chen 等ICSE 2023 · 被引用 71 次
- CCTEST: Testing and Repairing Code Completion SystemsZongjie Li, Chaozheng Wang, Zhibo Liu, Haoxuan Wang 等ICSE 2023 · 被引用 49 次
- Towards Understanding the Characteristics of Code Generation Errors Made by Large Language ModelsZhijie Wang, Zijie Zhou, Da Song, Yuheng Huang 等ICSE 2025 · 被引用 12 次
- AI Coders Are among Us: Rethinking Programming Language Grammar towards Efficient Code GenerationZhensu Sun, Xiaoning Du, Zhou Yang, Li Li 等ISSTA 2024 · 被引用 8 次
它引用的顶会 Paper11
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 被引用 1,224 次
- Q-BERT: Hessian Based Ultra Low Precision Quantization of BERTSheng Shen, Zhen Dong, Jiayu Ye, Linjian Ma 等AAAI 2020 · 被引用 656 次
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 被引用 541 次
- NatGen: generative pre-training by "naturalizing" source codeSaikat Chakraborty, Toufique Ahmed, Yangruibo Ding, Premkumar T. Devanbu 等FSE 2022 · 被引用 101 次
相关 Paper
- Natural Is the Best: Model-Agnostic Code Simplification for Pre-trained Large Language ModelsYan Wang, Xiaoning Li, Tien N. Nguyen, Shaohua Wang 等FSE 2024 · 被引用 6 次
- Towards Efficient Fine-Tuning of Pre-trained Code Models: An Experimental Study and BeyondEnsheng Shi, Yanlin Wang, Hongyu Zhang, Lun Du 等ISSTA 2023 · 被引用 37 次
- Compressing Pre-trained Models of Code into 3 MBJieke Shi, Zhou Yang, Bowen Xu, Hong Jin Kang 等ASE 2022 · 被引用 42 次
- GraphCodeBERT: Pre-training Code Representations with Data FlowDaya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng 等ICLR 2021 · 被引用 1,644 次
- LEANCODE: Understanding Models Better for Code Simplification of Pre-trained Large Language ModelsYan Wang, Ling Ding, Tien N. Nguyen, Shaohua Wang 等ACL 2025 · 被引用 1 次
