Learning semantic program embeddings with graph interval neural network
Yu Wang, Ke Wang, Fengjuan Gao, Linzhang Wang
摘要
Learning distributed representations of source code has been a challenging task for machine learning models. Earlier works treated programs as text so that natural language methods can be readily applied. Unfortunately, such approaches do not capitalize on the rich structural information possessed by source code. Of late, Graph Neural Network (GNN) was proposed to learn embeddings of programs from their graph representations. Due to the homogeneous (i.e. do not take advantage of the program-specific graph characteristics) and expensive (i.e. require heavy information exchange among nodes in the graph) message-passing procedure, GNN can suffer from precision issues, especially when dealing with programs rendered into large graphs. In this paper, we present a new graph neural architecture, called Graph Interval Neural Network (GINN), to tackle the weaknesses of the existing GNN. Unlike the standard GNN, GINN generalizes from a curated graph representation obtained through an abstraction method designed to aid models to learn. In particular, GINN focuses exclusively on intervals (generally manifested in looping construct) for mining the feature representation of a program, furthermore, GINN operates on a hierarchy of intervals for scaling the learning to large graphs. We evaluate GINN for two popular downstream applications: variable misuse prediction and method name prediction. Results show in both cases GINN outperforms the state-of-the-art models by a comfortable margin. We have also created a neural bug detector based on GINN to catch null pointer deference bugs in Java code. While learning from the same 9,000 methods extracted from 64 projects, GINN-based bug detector significantly outperforms GNN-based bug detector on 13 unseen test projects. Next, we deploy our trained GINN-based bug detector and Facebook Infer, arguably the state-of-the-art static analysis tool, to scan the codebase of 20 highly starred projects on GitHub. Through our manual inspection, we confirm 38 bugs out of 102 warnings raised by GINN-based bug detector compared to 34 bugs out of 129 warnings for Facebook Infer. We have reported 38 bugs GINN caught to developers, among which 11 have been fixed and 12 have been confirmed (fix pending). GINN has shown to be a general, powerful deep neural network for learning precise, semantic program embeddings.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper18
- CURE: Code-Aware Neural Machine Translation for Automatic Program RepairNan Jiang, Thibaud Lutellier, Lin TanICSE 2021 · 被引用 267 次
- Fully automated functional fuzzing of Android apps for detecting non-crashing logic bugsTing Su, Yichen Yan, Jue Wang, Jingling Sun 等OOPSLA 2021 · 被引用 58 次
- Lightweight global and local contexts guided method name recommendation with prior knowledgeShangwen Wang, Ming Wen, Bo Lin, Xiaoguang MaoFSE 2021 · 被引用 38 次
- Finding the dwarf: recovering precise types from WebAssembly binariesDaniel Lehmann, Michael PradelPLDI 2022 · 被引用 29 次
- Learning Graph-based Code Representations for Source-level Functional Similarity DetectionJiahao Liu, Jun Zeng, Xiang Wang, Zhenkai LiangICSE 2023 · 被引用 27 次
它引用的顶会 Paper3
- Hoppity: Learning Graph Transformations to Detect and Fix Bugs in ProgramsElizabeth Dinella, Hanjun Dai, Ziyang Li, Mayur Naik 等ICLR 2020 · 被引用 212 次
- LambdaNet: Probabilistic Type Inference using Graph Neural NetworksJiayi Wei, Maruth Goyal, Greg Durrett, Isil DilligICLR 2020 · 被引用 119 次
- Blended, precise semantic program embeddingsKe Wang, Zhendong SuPLDI 2020 · 被引用 53 次
相关 Paper
- Detecting Condition-Related Bugs with Control Flow Graph Neural NetworkJian Zhang, Xu Wang, Hongyu Zhang, Hailong Sun 等ISSTA 2023 · 被引用 19 次
- Detecting numerical bugs in neural network architecturesYuhao Zhang, Luyao Ren, Liqian Chen, Yingfei Xiong 等FSE 2020 · 被引用 66 次
- Path-sensitive code embedding via contrastive learning for software vulnerability detectionXiao Cheng, Guanqin Zhang, Haoyu Wang, Yulei SuiISSTA 2022 · 被引用 98 次
- Vulnerability Detection with Graph Simplification and Enhanced Graph Representation LearningXin-Cheng Wen, Yupan Chen, Cuiyun Gao, Hongyu Zhang 等ICSE 2023 · 被引用 77 次
- Global Relational Models of Source CodeVincent J. Hellendoorn, Charles Sutton, Rishabh Singh, Petros Maniatis 等ICLR 2020 · 被引用 252 次
