Keeping Pace with Ever-Increasing Data: Towards Continual Learning of Code Intelligence Models
Shuzheng Gao, Hongyu Zhang, Cuiyun Gao, Chaozheng Wang
Abstract
Previous research on code intelligence usually trains a deep learning model on a fixed dataset in an offline manner. However, in real-world scenarios, new code repositories emerge incessantly, and the carried new knowledge is beneficial for providing up-to-date code intelligence services to developers. In this paper, we aim at the following problem: How to enable code intelligence models to continually learn from ever-increasing data? One major challenge here is catastrophic forgetting, meaning that the model can easily forget knowledge learned from previous datasets when learning from the new dataset. To tackle this challenge, we propose REPEAT, a novel method for continual learning of code intelligence models. Specifically, REPEAT addresses the catastrophic forgetting problem with representative exemplars replay and adaptive parameter regularization. The representative exemplars replay component selects informative and diverse exemplars in each dataset and uses them to re-train model periodically. The adaptive parameter regularization component recognizes important parameters in the model and adaptively penalizes their changes to preserve the knowledge learned before. We evaluate the proposed approach on three code intelligence tasks including code summarization, software vulnerability detection, and code clone detection. Extensive experiments demonstrate that REPEAT consistently outperforms baseline methods on all tasks. For example, REPEAT improves the conventional fine-tuning method by 1.22, 5.61, and 1.72 on code summarization, vulnerability detection and clone detection, respectively.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a07307ac-b941-4f43-9bdd-d20e7e2e39baCited by top-tier papers8
- Evaluating and Improving ChatGPT for Unit Test GenerationZhiqiang Yuan, Mingwei Liu, Shiji Ding, Kaixin Wang et al.FSE 2024 · 89 citations
- What Makes Good In-Context Demonstrations for Code Intelligence Tasks with LLMs?Shuzheng Gao, Xin-Cheng Wen, Cuiyun Gao, Wenxuan Wang et al.ASE 2023 · 80 citations
- Two Birds with One Stone: Boosting Code Generation and Code Search via a Generative Adversarial NetworkShangwen Wang, Bo Lin, Zhensu Sun, Ming Wen et al.OOPSLA 2023 · 21 citations
- Learning in the Wild: Towards Leveraging Unlabeled Data for Effectively Tuning Pre-trained Code ModelsShuzheng Gao, Wenxin Mao, Cuiyun Gao, Li Li et al.ICSE 2024 · 15 citations
- CoSec: On-the-Fly Security Hardening of Code LLMs via Supervised Co-decodingDong Li, Meng Yan, Yaosheng Zhang, Zhongxin Liu et al.ISSTA 2024 · 10 citations
Builds on17
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- GraphCodeBERT: Pre-training Code Representations with Data FlowDaya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng et al.ICLR 2021 · 1,644 citations
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 1,224 citations
- O2U-Net: A Simple Noisy Label Detection Approach for Deep Neural NetworksJinchi Huang, Lie Qu, Rongfei Jia, Binqiang ZhaoICCV 2019 · 276 citations
- LAMOL: LAnguage MOdeling for Lifelong Language LearningFan-Keng Sun, Cheng-Hao Ho, Hung-Yi LeeICLR 2020 · 247 citations
Related papers
- On the Usage of Continual Learning for Out-of-Distribution Generalization in Pre-trained Language Models of CodeMartin Weyssow, Xin Zhou, Kisub Kim, David Lo et al.FSE 2023 · 9 citations
- No more fine-tuning? an experimental evaluation of prompt tuning in code intelligenceChaozheng Wang, Yuanhang Yang, Cuiyun Gao, Yun Peng et al.FSE 2022 · 148 citations
- Learning without Forgetting: Towards Continual learning of Fault Localization Models in Industrial Software SystemsChun Li, Hui Li, Zhong Li, Minxue Pan et al.ICSE 2026
- IDER: IDempotent Experience Replay for Reliable Continual LearningZhanwang Liu, Yuting Li, Haoyuan Gao, Yexin Li et al.ICLR 2026 · 5 citations
- CCT5: A Code-Change-Oriented Pre-trained ModelBo Lin, Shangwen Wang, Zhongxin Liu, Yepang Liu et al.FSE 2023 · 69 citations
