DeMinify: Neural Variable Name Recovery and Type Inference
Yi Li, Aashish Yadavally, Jiaxing Zhang, Shaohua Wang, Tien N. Nguyen
摘要
To avoid the exposure of original source code, the variable names deployed in the wild are often replaced by short, meaningless names, thus making the code difficult to understand and be analyzed. We introduce DeMinify, a Deep-Learning (DL)-based approach that formulates such recovery problem as the prediction of missing features in a Graph Convolutional Network–Missing Features. The graph represents both the relations among the variables and the relations among their types, in which the names or types of some nodes are missing. Moreover, DeMinify leverages dual-task learning to propagate the mutual impact between the learning of the variable names and that of their types. We conducted experiments to evaluate DeMinify in both name recovery and type prediction on a Python dataset with 180k methods and a JavaScript (JS) dataset with 322k files. For variable name prediction, in 76.7% and 81.6% of the cases in Python and JS code respectively, DeMinify can predict correctly the variables' names with a single suggested name. DeMinify relatively improves 15.3%–40.7% and 7.7%–49.7% in top-1 accuracy over the state-of-the-art variable name recovery approaches for Python and JS code, respectively. It also relatively improves 14.5%–51.9% in top-1 accuracy over the existing type prediction approaches. Our experimental results showed that learning of data types helps improve variable name recovery and vice versa.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- SelfPiCo: Self-Guided Partial Code Execution with LLMsZhipeng Xue, Zhipeng Gao, Shaohua Wang, Xing Hu 等ISSTA 2024 · 被引用 8 次
- Large Language Model-Aided Partial Program Dependence AnalysisXiaokai Rong, Aashish Yadavally, Tien N. NguyenICSE 2026 · 被引用 1 次
- DeCoMa: Detecting and Purifying Code Dataset Watermarks through Dual Channel Code AbstractionYuan Xiao, Yuchen Chen, Shiqing Ma, Haocheng Huang 等ISSTA 2025
- A New Approach to Evaluating Nullability Inference ToolsNima Karimipour, Erfan Arvan, Martin Kellogg, Manu SridharanFSE 2025
- JsDeObsBench: Measuring and Benchmarking LLMs for JavaScript DeobfuscationGuoqiang Chen, Xin Jin, Zhiqiang LinCCS 2025
它引用的顶会 Paper5
- LambdaNet: Probabilistic Type Inference using Graph Neural NetworksJiayi Wei, Maruth Goyal, Greg Durrett, Isil DilligICLR 2020 · 被引用 119 次
- TypeWriter: neural type prediction with search-based validationMichael Pradel, Georgios Gousios, Jason Liu, Satish ChandraFSE 2020 · 被引用 102 次
- Typilus: neural type hintsMiltiadis Allamanis, Earl T. Barr, Soline Ducousso, Zheng GaoPLDI 2020 · 被引用 92 次
- Type4Py: Practical Deep Similarity Learning-Based Type Inference for PythonAmir M. Mir, Evaldas Latoskinas, Sebastian Proksch, Georgios GousiosICSE 2022 · 被引用 59 次
- Static Inference Meets Deep learning: A Hybrid Type Inference Approach for PythonYun Peng, Cuiyun Gao, Zongjie Li, Bowei Gao 等ICSE 2022 · 被引用 48 次
相关 Paper
- Large-Scale and Language-Oblivious Code Authorship IdentificationMohammed Abuhamad, Tamer AbuHmed, Aziz Mohaisen, DaeHun NyangCCS 2018 · 被引用 102 次
- Nalin: learning from Runtime Behavior to Find Name-Value Inconsistencies in Jupyter NotebooksJibesh Patra, Michael PradelICSE 2022 · 被引用 14 次
- A Context-based Automated Approach for Method Name Consistency Checking and SuggestionYi Li, Shaohua Wang, Tien N. NguyenICSE 2021 · 被引用 36 次
- DLInfer: Deep Learning with Static Slicing for Python Type InferenceYanyan Yan, Yang Feng, Hongcheng Fan, Baowen XuICSE 2023 · 被引用 9 次
- Augmenting Decompiler Output with Learned Variable Names and TypesQibin Chen, Jeremy Lacomis, Edward J. Schwartz, Claire Le Goues 等USENIX Security 2022
