Learning to find naming issues with big code and small supervision
Jingxuan He, Cheng-Chun Lee, Veselin Raychev, Martin T. Vechev
摘要
We introduce a new approach for finding and fixing naming issues in source code. The method is based on a careful combination of unsupervised and supervised procedures: (i) unsupervised mining of patterns from Big Code that express common naming idioms. Program fragments violating such idioms indicates likely naming issues, and (ii) supervised learning of a classifier on a small labeled dataset which filters potential false positives from the violations.
We implemented our method in a system called Namer and evaluated it on a large number of Python and Java programs. We demonstrate that Namer is effective in finding naming mistakes in real world repositories with high precision (∼70%). Perhaps surprisingly, we also show that existing deep learning methods are not practically effective and achieve low precision in finding naming issues (up to ∼16%).
• Software and its engineering → Software defect analysis; • Theory of computation → Program analysis.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Self-Supervised Contrastive Learning for Code Retrieval and Summarization via Semantic-Preserving TransformationsNghi D. Q. Bui, Yijun Yu, Lingxiao JiangSIGIR 2021 · 被引用 98 次
- Path-sensitive code embedding via contrastive learning for software vulnerability detectionXiao Cheng, Guanqin Zhang, Haoyu Wang, Yulei SuiISSTA 2022 · 被引用 98 次
- On Distribution Shift in Learning-based Bug DetectorsJingxuan He, Luca Beurer-Kellner, Martin T. VechevICML 2022 · 被引用 20 次
- Nalin: learning from Runtime Behavior to Find Name-Value Inconsistencies in Jupyter NotebooksJibesh Patra, Michael PradelICSE 2022 · 被引用 14 次
- DAInfer: Inferring API Aliasing Specifications from Library Documentation via Neurosymbolic OptimizationChengpeng Wang, Jipeng Zhang, Rongxin Wu, Charles ZhangFSE 2024 · 被引用 5 次
它引用的顶会 Paper9
- Coverage-based Greybox Fuzzing as Markov ChainMarcel Böhme, Van-Thuan Pham, Abhik RoychoudhuryCCS 2016 · 被引用 1,026 次
- Skyfire: Data-Driven Seed Generation for FuzzingJunjie Wang, Bihuan Chen, Lei Wei, Yang LiuS&P 2017 · 被引用 382 次
- Learning to Fuzz from Symbolic Execution with Application to Smart ContractsJingxuan He, Mislav Balunovic, Nodar Ambroladze, Petar Tsankov 等CCS 2019 · 被引用 288 次
- Global Relational Models of Source CodeVincent J. Hellendoorn, Charles Sutton, Rishabh Singh, Petros Maniatis 等ICLR 2020 · 被引用 252 次
- NEUZZ: Efficient Fuzzing with Neural Program SmoothingDongdong She, Kexin Pei, Dave Epstein, Junfeng Yang 等S&P 2019 · 被引用 220 次
相关 Paper
- A Context-based Automated Approach for Method Name Consistency Checking and SuggestionYi Li, Shaohua Wang, Tien N. NguyenICSE 2021 · 被引用 36 次
- Streamlining Java Programming: Uncovering Well-Formed Idioms with IdioMineYanming Yang, Xing Hu, Xin Xia, David Lo 等ICSE 2024 · 被引用 2 次
- Big code != big vocabulary: open-vocabulary models for source codeRafael-Michael Karampatsis, Hlib Babii, Romain Robbes, Charles Sutton 等ICSE 2020 · 被引用 140 次
- Suggesting natural method names to check name consistenciesSon Nguyen, Hung Phan, Trinh Le, Tien N. NguyenICSE 2020 · 被引用 64 次
- Learning to Recommend Method Names with Global ContextFang Liu, Ge Li, Zhiyi Fu, Shuai Lu 等ICSE 2022 · 被引用 32 次
