Generating Variable Explanations via Zero-shot Prompt Learning
Chong Wang, Yiling Lou, Junwei Liu, Xin Peng
Abstract
As basic elements in program, variables convey essential information that is critical for program comprehension and maintenance. However, understanding the meanings of variables in program is not always easy for developers, since poor-quality variable names are prevalent while such variable are less informative for program comprehension. Therefore, in this paper, we target at generating concise natural language explanations for variables to facilitate program comprehension. In particular, there are two challenges in variable explanation generation, including the lack of training data and the association with complex code contexts around the variable. To address these issues, we propose a novel approach ZeroVar,which leverages code pre-trained models and zero-shot prompt learning to generate explanations for the variable based on its code context. ZeroVarcontains two stages: (i) a pre-training stage that continually pre-trains a base model (i.e., CodeT5) to recover the randomly-masked parameter descriptions in method docstrings; and (ii) a zero-shot prompt learning stage that leverages the pre-trained model to generate explanations for a given variable via the prompt constructed with the variable and its belonging method context. We then extensively evaluate the quality and usefulness of the variable explanations generated by ZeroVar.We construct an evaluation dataset of 773 variables and their reference explanations. Our results show that ZeroVarcan generate higher-quality explanations than baselines, not only on automated metrics such as BLEU and ROUGE, but also on human metrics such as correctness, completeness, and conciseness. Moreover, we further assess the usefulness of ZeroVAR-generated explanations on two downstream tasks related to variable naming quality, i.e., abbreviation expansion and spelling correction. For abbreviation expansion, the generated variable explanations can help improve the present rate (+13.1%), precision (+3.6%), and recall (+10.0%) of the state-of-the-art abbreviation explanation approach. For spelling correction, by using the generated explanations we can achieve higher hit@1 (+162.9(%) and hit@3 (+49.6%) than the recent variable representation learning approach.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 5e39ade9-2d57-43c8-ba52-51baa772db1eCited by top-tier papers4
- LLMDFA: Analyzing Dataflow in Code with Large Language ModelsChengpeng Wang, Wuqi Zhang, Zian Su, Xiangzhe Xu et al.NeurIPS 2024 · 51 citations
- Smart-LLaMA-DPO: Reinforced Large Language Model for Explainable Smart Contract Vulnerability DetectionLei Yu, Zhirong Huang, Hang Yuan, Shiqi Cheng et al.ISSTA 2025 · 13 citations
- TIGER: A Generating-Then-Ranking Framework for Practical Python Type InferenceChong Wang, Jian Zhang, Yiling Lou, Mingwei Liu et al.ICSE 2025 · 1 citation
- PriAgent: A Collaborative Multi-Agent Framework for Auditing Android Privacy ComplianceZiwei Zhang, Zhao Li, Zhuojun Jiang, Jiangyi Yin et al.AAAI 2026
Related papers
- VarCLR: Variable Semantic Representation Pre-training via Contrastive LearningQibin Chen, Jeremy Lacomis, Edward J. Schwartz, Graham Neubig et al.ICSE 2022 · 23 citations
- DocPrompting: Generating Code by Retrieving the DocsShuyan Zhou, Uri Alon, Frank F. Xu, Zhengbao Jiang et al.ICLR 2023 · 18 citations
- No more fine-tuning? an experimental evaluation of prompt tuning in code intelligenceChaozheng Wang, Yuanhang Yang, Cuiyun Gao, Yun Peng et al.FSE 2022 · 148 citations
- Retrieval-Based Prompt Selection for Code-Related Few-Shot LearningNoor Nashid, Mifta Sintaha, Ali MesbahICSE 2023 · 156 citations
- RefBERT: A Two-Stage Pre-trained Framework for Automatic Rename RefactoringHao Liu, Yanlin Wang, Zhao Wei, Yong Xu et al.ISSTA 2023 · 19 citations
