Learning to find naming issues with big code and small supervision
Jingxuan He, Cheng-Chun Lee, Veselin Raychev, Martin T. Vechev
Abstract
We introduce a new approach for finding and fixing naming issues in source code. The method is based on a careful combination of unsupervised and supervised procedures: (i) unsupervised mining of patterns from Big Code that express common naming idioms. Program fragments violating such idioms indicates likely naming issues, and (ii) supervised learning of a classifier on a small labeled dataset which filters potential false positives from the violations.
We implemented our method in a system called Namer and evaluated it on a large number of Python and Java programs. We demonstrate that Namer is effective in finding naming mistakes in real world repositories with high precision (∼70%). Perhaps surprisingly, we also show that existing deep learning methods are not practically effective and achieve low precision in finding naming issues (up to ∼16%).
• Software and its engineering → Software defect analysis; • Theory of computation → Program analysis.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 680fe88c-5f6f-4e99-a6b6-94977f446603Cited by top-tier papers5
- Self-Supervised Contrastive Learning for Code Retrieval and Summarization via Semantic-Preserving TransformationsNghi D. Q. Bui, Yijun Yu, Lingxiao JiangSIGIR 2021 · 98 citations
- Path-sensitive code embedding via contrastive learning for software vulnerability detectionXiao Cheng, Guanqin Zhang, Haoyu Wang, Yulei SuiISSTA 2022 · 98 citations
- On Distribution Shift in Learning-based Bug DetectorsJingxuan He, Luca Beurer-Kellner, Martin T. VechevICML 2022 · 20 citations
- Nalin: learning from Runtime Behavior to Find Name-Value Inconsistencies in Jupyter NotebooksJibesh Patra, Michael PradelICSE 2022 · 14 citations
- DAInfer: Inferring API Aliasing Specifications from Library Documentation via Neurosymbolic OptimizationChengpeng Wang, Jipeng Zhang, Rongxin Wu, Charles ZhangFSE 2024 · 5 citations
Builds on9
- Coverage-based Greybox Fuzzing as Markov ChainMarcel Böhme, Van-Thuan Pham, Abhik RoychoudhuryCCS 2016 · 1,026 citations
- Skyfire: Data-Driven Seed Generation for FuzzingJunjie Wang, Bihuan Chen, Lei Wei, Yang LiuS&P 2017 · 382 citations
- Learning to Fuzz from Symbolic Execution with Application to Smart ContractsJingxuan He, Mislav Balunovic, Nodar Ambroladze, Petar Tsankov et al.CCS 2019 · 288 citations
- Global Relational Models of Source CodeVincent J. Hellendoorn, Charles Sutton, Rishabh Singh, Petros Maniatis et al.ICLR 2020 · 252 citations
- NEUZZ: Efficient Fuzzing with Neural Program SmoothingDongdong She, Kexin Pei, Dave Epstein, Junfeng Yang et al.S&P 2019 · 220 citations
Related papers
- A Context-based Automated Approach for Method Name Consistency Checking and SuggestionYi Li, Shaohua Wang, Tien N. NguyenICSE 2021 · 36 citations
- Streamlining Java Programming: Uncovering Well-Formed Idioms with IdioMineYanming Yang, Xing Hu, Xin Xia, David Lo et al.ICSE 2024 · 2 citations
- Big code != big vocabulary: open-vocabulary models for source codeRafael-Michael Karampatsis, Hlib Babii, Romain Robbes, Charles Sutton et al.ICSE 2020 · 140 citations
- Suggesting natural method names to check name consistenciesSon Nguyen, Hung Phan, Trinh Le, Tien N. NguyenICSE 2020 · 64 citations
- Learning to Recommend Method Names with Global ContextFang Liu, Ge Li, Zhiyi Fu, Shuai Lu et al.ICSE 2022 · 32 citations
