Nalin: learning from Runtime Behavior to Find Name-Value Inconsistencies in Jupyter Notebooks
Jibesh Patra, Michael Pradel
Abstract
Variable names are important to understand and maintain code. If a variable name and the value stored in the variable do not match, then the program suffers from a name-value inconsistency, which is due to one of two situations that developers may want to fix: Either a correct value is referred to through a misleading name, which negatively affects code understandability and maintainability, or the correct name is bound to a wrong value, which may cause unexpected runtime behavior. Finding name-value inconsistencies is hard because it requires an understanding of the meaning of names and knowledge about the values assigned to a variable at runtime. This paper presents Nalin, a technique to automatically detect name-value inconsistencies. The approach combines a dynamic analysis that tracks assignments of values to names with a neural machine learning model that predicts whether a name and a value fit together. To the best of our knowledge, this is the first work to formulate the problem of finding coding issues as a classification problem over names and runtime values. We apply Nalin to 106,652 real-world Python programs, where meaningful names are particularly important due to the absence of statically declared types. Our results show that the classifier detects name-value inconsistencies with high accuracy, that the warnings reported by Nalin have a precision of 80% and a recall of 76% w.r.t. a ground truth created in a user study, and that our approach complements existing techniques for finding coding issues.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext db53779f-ab96-44be-8d6b-8fa3a88af95bCited by top-tier papers6
- The evolution of type annotations in python: an empirical studyLuca Di Grazia, Michael PradelFSE 2022 · 29 citations
- TRACED: Execution-aware Pre-training for Source CodeYangruibo Ding, Benjamin Steenhoek, Kexin Pei, Gail E. Kaiser et al.ICSE 2024 · 29 citations
- LExecutor: Learning-Guided ExecutionBeatriz Souza, Michael PradelFSE 2023 · 16 citations
- Beware of the Unexpected: Bimodal Taint AnalysisYiu Wai Chow, Max Schäfer, Michael PradelISSTA 2023 · 13 citations
- NeuDep: neural binary memory dependence analysisKexin Pei, Dongdong She, Michael Wang, Scott Geng et al.FSE 2022 · 8 citations
Builds on16
- Global Relational Models of Source CodeVincent J. Hellendoorn, Charles Sutton, Rishabh Singh, Petros Maniatis et al.ICLR 2020 · 252 citations
- Hoppity: Learning Graph Transformations to Detect and Fix Bugs in ProgramsElizabeth Dinella, Hanjun Dai, Ziyang Li, Mayur Naik et al.ICLR 2020 · 212 citations
- Code Prediction by Feeding Trees to TransformersSeohyun Kim, Jinman Zhao, Yuchi Tian, Satish ChandraICSE 2021 · 179 citations
- CC2Vec: distributed representations of code changesThong Hoang, Hong Jin Kang, David Lo, Julia LawallICSE 2020 · 169 citations
- LambdaNet: Probabilistic Type Inference using Graph Neural NetworksJiayi Wei, Maruth Goyal, Greg Durrett, Isil DilligICLR 2020 · 119 citations
Related papers
- When to Say What: Learning to Find Condition-Message InconsistenciesIslem Bouzenia, Michael PradelICSE 2023 · 5 citations
- DyLin: A Dynamic Linter for PythonAryaz Eghbali, Felix Burk, Michael PradelFSE 2025 · 1 citation
- A Context-based Automated Approach for Method Name Consistency Checking and SuggestionYi Li, Shaohua Wang, Tien N. NguyenICSE 2021 · 36 citations
- Static Type Recommendation for PythonKe Sun, Yifan Zhao, Dan Hao, Lu ZhangASE 2022 · 6 citations
- DeMinify: Neural Variable Name Recovery and Type InferenceYi Li, Aashish Yadavally, Jiaxing Zhang, Shaohua Wang et al.FSE 2023 · 6 citations
