On-the-Fly Adapting Code Summarization on Trainable Cost-Effective Language Models
Yufan Cai, Yun Lin, Chenyan Liu, Jinglian Wu, Yifan Zhang, Yiming Liu, Yeyun Gong, Jin Song Dong
摘要
Deep learning models are emerging to summarize source code to comment for code documentation and program comprehension. We can achieve good performance by training the model on large training corpus. However, in practice, the code samples from different projects can have contradictory training signal for learning a deep comment generator, making the model struggled to fit all the training samples. In this work, we introduce a novel approach, AdaCom, to improve the performance of comment generators by on-the-fly model adaptation. This research is motivated by the observation that deep comment generators often need to strike a balance as they need to fit all the training samples. Specifically, for one certain target code c, some training samples S p could have made more contributions while other samples S o could have counter effects. However, the traditional fine-tuned models need to fit both S p and S o from a global perspective, leading to compromised performance for one certain target code c. In this context, we design AdaCom to (1) detect whether the model might have a compromised performance on a target code c and (2) retrieve a few helpful training samples S p that have contradictory samples in the training dataset and, (3) adapt the model on the fly by re-training the S p to strengthen the helpful samples and unlearn the harmful samples. Our extensive experiments on 7 comment generators and 4 public datasets show that (1) AdaCom can significantly boost the performance of comment generation (BLEU4 score by on average 14.9%, METEOR by 12.2%, and ROUGE-L by 7.4%), ( 2 ) the adaptation on one code sample is cost-effective and acceptable as an on-the-fly solution, and (3) AdaCom can adapt well on out-of-distribution code samples.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Towards Understanding the Characteristics of Code Generation Errors Made by Large Language ModelsZhijie Wang, Zijie Zhou, Da Song, Yuheng Huang 等ICSE 2025 · 被引用 12 次
- CoEdPilot: Recommending Code Edits with Learned Prior Edit Relevance, Project-wise Awareness, and Interactive NatureChenyan Liu, Yufan Cai, Yun Lin, Yuhuan Huang 等ISSTA 2024 · 被引用 7 次
- REACCEPT: Automated Co-evolution of Production and Test Code Based on Dynamic Validation and Large Language ModelsJianlei Chi, Xiaotian Wang, Yuhan Huang, Lechen Yu 等ISSTA 2025 · 被引用 2 次
- Learning Project-wise Subsequent Code Edits via Interleaving Neural-based Induction and Tool-based DeductionChenyan Liu, Yun Lin, Yuhuan Huang, Jiaxin Chang 等ASE 2025 · 被引用 1 次
- EditFlow: Benchmarking and Optimizing Code Edit Recommendation Systems via Reconstruction of Developer FlowsChenyan Liu, Yun Lin, Jiaxin Chang, Jiawei Liu 等OOPSLA 2026
它引用的顶会 Paper12
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Tent: Fully Test-Time Adaptation by Entropy MinimizationDequan Wang, Evan Shelhamer, Shaoteng Liu, Bruno A. Olshausen 等ICLR 2021 · 被引用 1,731 次
- GraphCodeBERT: Pre-training Code Representations with Data FlowDaya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng 等ICLR 2021 · 被引用 1,644 次
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 被引用 1,224 次
- Test-Time Training with Self-Supervision for Generalization under Distribution ShiftsYu Sun, Xiaolong Wang, Zhuang Liu, John Miller 等ICML 2020 · 被引用 1,220 次
相关 Paper
- Code to Comment "Translation": Data, Metrics, Baselining & EvaluationDavid Gros, Hariharan Sezhiyan, Prem Devanbu, Zhou YuASE 2020 · 被引用 43 次
- On the Evaluation of Neural Code SummarizationEnsheng Shi, Yanlin Wang, Lun Du, Junjie Chen 等ICSE 2022 · 被引用 76 次
- Deep Just-In-Time Inconsistency Detection Between Comments and Source CodeSheena Panthaplackel, Junyi Jessy Li, Milos Gligoric, Raymond J. MooneyAAAI 2021 · 被引用 62 次
- Retrieve and Refine: Exemplar-based Neural Comment GenerationBolin Wei, Yongmin Li, Ge Li, Xin Xia 等ASE 2020 · 被引用 68 次
- Code Difference Guided Adversarial Example Generation for Deep Code ModelsZhao Tian, Junjie Chen, Zhi JinASE 2023 · 被引用 27 次
