VarCLR: Variable Semantic Representation Pre-training via Contrastive Learning
Qibin Chen, Jeremy Lacomis, Edward J. Schwartz, Graham Neubig, Bogdan Vasilescu, Claire Le Goues
摘要
Variable names are critical for conveying intended program behavior. Machine learning-based program analysis methods use variable name representations for a wide range of tasks, such as suggesting new variable names and bug detection. Ideally, such methods could capture semantic relationships between names beyond syntactic similarity, e.g., the fact that the names average and mean are similar. Unfortunately, previous work has found that even the best of previous representation approaches primarily capture "relatedness" (whether two variables are linked at all), rather than "similarity" (whether they actually have the same meaning). We propose VarCLR, a new approach for learning semantic representations of variable names that effectively captures variable similarity in this stricter sense. We observe that this problem is an excellent fit for contrastive learning, which aims to minimize the distance between explicitly similar inputs, while maximizing the distance between dissimilar inputs. This requires labeled training data, and thus we construct a novel, weakly-supervised variable renaming dataset mined from GitHub edits. We show that VarCLR enables the effective application of sophisticated, general-purpose language models like BERT, to variable name representation and thus also to related downstream tasks like variable name similarity search or spelling correction. VarCLR produces models that significantly outperform the state-of-the-art on IdBench, an existing benchmark that explicitly captures variable similarity (as distinct from relatedness). Finally, we contribute a release of all data, code, and pre-trained models, aiming to provide a drop-in replacement for variable representations used in either existing or future program analyses that rely on variable names. INTRODUCTION Variable names convey key information about code structure and developer intention. They are thus central for code comprehension, readability, and maintainability [7, 46]. A growing array of
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- CodeRL: Mastering Code Generation through Pretrained Models and Deep Reinforcement LearningHung Le, Yue Wang, Akhilesh Deepak Gotmare, Silvio Savarese 等NeurIPS 2022 · 被引用 571 次
- Demo2Code: From Summarizing Demonstrations to Synthesizing Code via Extended Chain-of-ThoughtYuki Wang, Gonzalo Gonzalez-Pumariega, Yash Sharma, Sanjiban ChoudhuryNeurIPS 2023 · 被引用 67 次
- ContraBERT: Enhancing Code Pre-trained Models via Contrastive LearningShangqing Liu, Bozhi Wu, Xiaofei Xie, Guozhu Meng 等ICSE 2023 · 被引用 56 次
- Distinguishing Look-Alike Innocent and Vulnerable Code by Subtle Semantic Representation Learning and ExplanationChao Ni, Xin Yin, Kaiwen Yang, Dehai Zhao 等FSE 2023 · 被引用 42 次
- Ahoy SAILR! There is No Need to DREAM of C: A Compiler-Aware Structuring Algorithm for Binary DecompilationZion Leonahenahe Basque, Ati Priya Bajaj, Wil Gibbs, Jude O'Kain 等USENIX Security 2024 · 被引用 32 次
它引用的顶会 Paper16
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna 等NeurIPS 2020 · 被引用 7,049 次
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal 等NeurIPS 2020 · 被引用 5,249 次
相关 Paper
- IdBench: Evaluating Semantic Representations of Identifier Names in Source CodeYaza Wainakh, Moiz Rauf, Michael PradelICSE 2021 · 被引用 2 次
- "Len or index or count, anything but v1": Predicting Variable Names in Decompilation Output with Transfer LearningKuntal Kumar Pal, Ati Priya Bajaj, Pratyay Banerjee, Audrey Dutcher 等S&P 2024 · 被引用 30 次
- Generating Variable Explanations via Zero-shot Prompt LearningChong Wang, Yiling Lou, Junwei Liu, Xin PengASE 2023 · 被引用 7 次
- RefBERT: A Two-Stage Pre-trained Framework for Automatic Rename RefactoringHao Liu, Yanlin Wang, Zhao Wei, Yong Xu 等ISSTA 2023 · 被引用 19 次
- CLEAR: Contrastive Learning for API RecommendationMoshi Wei, Nima Shiri Harzevili, Yuchao Huang, Junjie Wang 等ICSE 2022 · 被引用 43 次
