Nalin: learning from Runtime Behavior to Find Name-Value Inconsistencies in Jupyter Notebooks
Jibesh Patra, Michael Pradel
摘要
Variable names are important to understand and maintain code. If a variable name and the value stored in the variable do not match, then the program suffers from a name-value inconsistency, which is due to one of two situations that developers may want to fix: Either a correct value is referred to through a misleading name, which negatively affects code understandability and maintainability, or the correct name is bound to a wrong value, which may cause unexpected runtime behavior. Finding name-value inconsistencies is hard because it requires an understanding of the meaning of names and knowledge about the values assigned to a variable at runtime. This paper presents Nalin, a technique to automatically detect name-value inconsistencies. The approach combines a dynamic analysis that tracks assignments of values to names with a neural machine learning model that predicts whether a name and a value fit together. To the best of our knowledge, this is the first work to formulate the problem of finding coding issues as a classification problem over names and runtime values. We apply Nalin to 106,652 real-world Python programs, where meaningful names are particularly important due to the absence of statically declared types. Our results show that the classifier detects name-value inconsistencies with high accuracy, that the warnings reported by Nalin have a precision of 80% and a recall of 76% w.r.t. a ground truth created in a user study, and that our approach complements existing techniques for finding coding issues.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- The evolution of type annotations in python: an empirical studyLuca Di Grazia, Michael PradelFSE 2022 · 被引用 29 次
- TRACED: Execution-aware Pre-training for Source CodeYangruibo Ding, Benjamin Steenhoek, Kexin Pei, Gail E. Kaiser 等ICSE 2024 · 被引用 29 次
- LExecutor: Learning-Guided ExecutionBeatriz Souza, Michael PradelFSE 2023 · 被引用 16 次
- Beware of the Unexpected: Bimodal Taint AnalysisYiu Wai Chow, Max Schäfer, Michael PradelISSTA 2023 · 被引用 13 次
- NeuDep: neural binary memory dependence analysisKexin Pei, Dongdong She, Michael Wang, Scott Geng 等FSE 2022 · 被引用 8 次
它引用的顶会 Paper16
- Global Relational Models of Source CodeVincent J. Hellendoorn, Charles Sutton, Rishabh Singh, Petros Maniatis 等ICLR 2020 · 被引用 252 次
- Hoppity: Learning Graph Transformations to Detect and Fix Bugs in ProgramsElizabeth Dinella, Hanjun Dai, Ziyang Li, Mayur Naik 等ICLR 2020 · 被引用 212 次
- Code Prediction by Feeding Trees to TransformersSeohyun Kim, Jinman Zhao, Yuchi Tian, Satish ChandraICSE 2021 · 被引用 179 次
- CC2Vec: distributed representations of code changesThong Hoang, Hong Jin Kang, David Lo, Julia LawallICSE 2020 · 被引用 169 次
- LambdaNet: Probabilistic Type Inference using Graph Neural NetworksJiayi Wei, Maruth Goyal, Greg Durrett, Isil DilligICLR 2020 · 被引用 119 次
相关 Paper
- When to Say What: Learning to Find Condition-Message InconsistenciesIslem Bouzenia, Michael PradelICSE 2023 · 被引用 5 次
- DyLin: A Dynamic Linter for PythonAryaz Eghbali, Felix Burk, Michael PradelFSE 2025 · 被引用 1 次
- A Context-based Automated Approach for Method Name Consistency Checking and SuggestionYi Li, Shaohua Wang, Tien N. NguyenICSE 2021 · 被引用 36 次
- Static Type Recommendation for PythonKe Sun, Yifan Zhao, Dan Hao, Lu ZhangASE 2022 · 被引用 6 次
- DeMinify: Neural Variable Name Recovery and Type InferenceYi Li, Aashish Yadavally, Jiaxing Zhang, Shaohua Wang 等FSE 2023 · 被引用 6 次
