Generating Variable Explanations via Zero-shot Prompt Learning
Chong Wang, Yiling Lou, Junwei Liu, Xin Peng
摘要
As basic elements in program, variables convey essential information that is critical for program comprehension and maintenance. However, understanding the meanings of variables in program is not always easy for developers, since poor-quality variable names are prevalent while such variable are less informative for program comprehension. Therefore, in this paper, we target at generating concise natural language explanations for variables to facilitate program comprehension. In particular, there are two challenges in variable explanation generation, including the lack of training data and the association with complex code contexts around the variable. To address these issues, we propose a novel approach ZeroVar,which leverages code pre-trained models and zero-shot prompt learning to generate explanations for the variable based on its code context. ZeroVarcontains two stages: (i) a pre-training stage that continually pre-trains a base model (i.e., CodeT5) to recover the randomly-masked parameter descriptions in method docstrings; and (ii) a zero-shot prompt learning stage that leverages the pre-trained model to generate explanations for a given variable via the prompt constructed with the variable and its belonging method context. We then extensively evaluate the quality and usefulness of the variable explanations generated by ZeroVar.We construct an evaluation dataset of 773 variables and their reference explanations. Our results show that ZeroVarcan generate higher-quality explanations than baselines, not only on automated metrics such as BLEU and ROUGE, but also on human metrics such as correctness, completeness, and conciseness. Moreover, we further assess the usefulness of ZeroVAR-generated explanations on two downstream tasks related to variable naming quality, i.e., abbreviation expansion and spelling correction. For abbreviation expansion, the generated variable explanations can help improve the present rate (+13.1%), precision (+3.6%), and recall (+10.0%) of the state-of-the-art abbreviation explanation approach. For spelling correction, by using the generated explanations we can achieve higher hit@1 (+162.9(%) and hit@3 (+49.6%) than the recent variable representation learning approach.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper4
- LLMDFA: Analyzing Dataflow in Code with Large Language ModelsChengpeng Wang, Wuqi Zhang, Zian Su, Xiangzhe Xu 等NeurIPS 2024 · 被引用 51 次
- Smart-LLaMA-DPO: Reinforced Large Language Model for Explainable Smart Contract Vulnerability DetectionLei Yu, Zhirong Huang, Hang Yuan, Shiqi Cheng 等ISSTA 2025 · 被引用 13 次
- TIGER: A Generating-Then-Ranking Framework for Practical Python Type InferenceChong Wang, Jian Zhang, Yiling Lou, Mingwei Liu 等ICSE 2025 · 被引用 1 次
- PriAgent: A Collaborative Multi-Agent Framework for Auditing Android Privacy ComplianceZiwei Zhang, Zhao Li, Zhuojun Jiang, Jiangyi Yin 等AAAI 2026
相关 Paper
- VarCLR: Variable Semantic Representation Pre-training via Contrastive LearningQibin Chen, Jeremy Lacomis, Edward J. Schwartz, Graham Neubig 等ICSE 2022 · 被引用 23 次
- DocPrompting: Generating Code by Retrieving the DocsShuyan Zhou, Uri Alon, Frank F. Xu, Zhengbao Jiang 等ICLR 2023 · 被引用 18 次
- No more fine-tuning? an experimental evaluation of prompt tuning in code intelligenceChaozheng Wang, Yuanhang Yang, Cuiyun Gao, Yun Peng 等FSE 2022 · 被引用 148 次
- Retrieval-Based Prompt Selection for Code-Related Few-Shot LearningNoor Nashid, Mifta Sintaha, Ali MesbahICSE 2023 · 被引用 156 次
- RefBERT: A Two-Stage Pre-trained Framework for Automatic Rename RefactoringHao Liu, Yanlin Wang, Zhao Wei, Yong Xu 等ISSTA 2023 · 被引用 19 次
