VarCLR: Variable Semantic Representation Pre-training via Contrastive Learning
Qibin Chen, Jeremy Lacomis, Edward J. Schwartz, Graham Neubig, Bogdan Vasilescu, Claire Le Goues
Abstract
Variable names are critical for conveying intended program behavior. Machine learning-based program analysis methods use variable name representations for a wide range of tasks, such as suggesting new variable names and bug detection. Ideally, such methods could capture semantic relationships between names beyond syntactic similarity, e.g., the fact that the names average and mean are similar. Unfortunately, previous work has found that even the best of previous representation approaches primarily capture "relatedness" (whether two variables are linked at all), rather than "similarity" (whether they actually have the same meaning). We propose VarCLR, a new approach for learning semantic representations of variable names that effectively captures variable similarity in this stricter sense. We observe that this problem is an excellent fit for contrastive learning, which aims to minimize the distance between explicitly similar inputs, while maximizing the distance between dissimilar inputs. This requires labeled training data, and thus we construct a novel, weakly-supervised variable renaming dataset mined from GitHub edits. We show that VarCLR enables the effective application of sophisticated, general-purpose language models like BERT, to variable name representation and thus also to related downstream tasks like variable name similarity search or spelling correction. VarCLR produces models that significantly outperform the state-of-the-art on IdBench, an existing benchmark that explicitly captures variable similarity (as distinct from relatedness). Finally, we contribute a release of all data, code, and pre-trained models, aiming to provide a drop-in replacement for variable representations used in either existing or future program analyses that rely on variable names. INTRODUCTION Variable names convey key information about code structure and developer intention. They are thus central for code comprehension, readability, and maintainability [7, 46]. A growing array of
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a65e717b-dae0-463c-ae96-d16113686a37Cited by top-tier papers13
- CodeRL: Mastering Code Generation through Pretrained Models and Deep Reinforcement LearningHung Le, Yue Wang, Akhilesh Deepak Gotmare, Silvio Savarese et al.NeurIPS 2022 · 571 citations
- Demo2Code: From Summarizing Demonstrations to Synthesizing Code via Extended Chain-of-ThoughtYuki Wang, Gonzalo Gonzalez-Pumariega, Yash Sharma, Sanjiban ChoudhuryNeurIPS 2023 · 67 citations
- ContraBERT: Enhancing Code Pre-trained Models via Contrastive LearningShangqing Liu, Bozhi Wu, Xiaofei Xie, Guozhu Meng et al.ICSE 2023 · 56 citations
- Distinguishing Look-Alike Innocent and Vulnerable Code by Subtle Semantic Representation Learning and ExplanationChao Ni, Xin Yin, Kaiwen Yang, Dehai Zhao et al.FSE 2023 · 42 citations
- Ahoy SAILR! There is No Need to DREAM of C: A Compiler-Aware Structuring Algorithm for Binary DecompilationZion Leonahenahe Basque, Ati Priya Bajaj, Wil Gibbs, Jude O'Kain et al.USENIX Security 2024 · 32 citations
Builds on16
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna et al.NeurIPS 2020 · 7,049 citations
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal et al.NeurIPS 2020 · 5,249 citations
Related papers
- IdBench: Evaluating Semantic Representations of Identifier Names in Source CodeYaza Wainakh, Moiz Rauf, Michael PradelICSE 2021 · 2 citations
- "Len or index or count, anything but v1": Predicting Variable Names in Decompilation Output with Transfer LearningKuntal Kumar Pal, Ati Priya Bajaj, Pratyay Banerjee, Audrey Dutcher et al.S&P 2024 · 30 citations
- Generating Variable Explanations via Zero-shot Prompt LearningChong Wang, Yiling Lou, Junwei Liu, Xin PengASE 2023 · 7 citations
- RefBERT: A Two-Stage Pre-trained Framework for Automatic Rename RefactoringHao Liu, Yanlin Wang, Zhao Wei, Yong Xu et al.ISSTA 2023 · 19 citations
- CLEAR: Contrastive Learning for API RecommendationMoshi Wei, Nima Shiri Harzevili, Yuchao Huang, Junjie Wang et al.ICSE 2022 · 43 citations
