CodeGrid: A Grid Representation of Code
Abdoul Kader Kaboré, Earl T. Barr, Jacques Klein, Tegawendé F. Bissyandé
摘要
Code representation is a key step in the application of AI in software engineering. Generic NLP representations are e ective but do not exploit all the rich structure inherent to code. Recent work has focused on extracting abstract syntax trees (AST) and integrating their structural information into code representations. These AST-enhanced representations advanced the state of the art and accelerated new applications of AI to software engineering. ASTs, however, neglect important aspects of code structure, notably control and data ow, leaving some potentially relevant code signal unexploited. For example, purely image-based representations perform nearly as well as AST-based representations, despite the fact that they must learn to even recognize tokens, let alone their semantics. This result, from prior work, is strong evidence that these new code representations can still be improved; it also raises the question of just what signal image-based approaches are exploiting. We answer this question. We show that code is spatial and exploit this fact to propose CodeGrid, a new representation that embeds tokens into a grid that preserves code layout. Unlike some of the existing state of the art, CodeGrid is agnostic to the downstream task: whether that task is generation or classi cation, CodeGrid can complement the learning algorithm with spatial signal. For example, we show that CNNs, which are inherently spatially-aware models, can exploit CodeGrid outputs to e ectively tackle fundamental software engineering tasks, such as code classi cation, code clone detection and vulnerability detection. PixelCNN leverages CodeGrid's grid representations to achieve code completion. Through extensive experiments, we validate our spatial code hypothesis, quantifying model performance as we vary the degree to which the representation preserves the grid. To demonstrate its generality, we show that CodeGrid augments models, improving their performance on a range of tasks. On clone detection, CodeGrid improves ASTNN's performance by 3.3% F1 score. * Some work carried out while a visiting scholar at Google DeepMind.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper5
- VUDDY: A Scalable Approach for Vulnerable Code Clone DiscoverySeulbae Kim, Seunghoon Woo, Heejo Lee, Hakjoo OhS&P 2017 · 被引用 388 次
- Multi-task Learning based Pre-trained Language Model for Code CompletionFang Liu, Ge Li, Yunfei Zhao, Zhi JinASE 2020 · 被引用 162 次
- Structural Language Models of CodeUri Alon, Roy Sadaka, Omer Levy, Eran YahavICML 2020 · 被引用 115 次
- Software visualization and deep transfer learning for effective software defect predictionJinyin Chen, Keke Hu, Yue Yu, Zhuangzhi Chen 等ICSE 2020 · 被引用 89 次
- VulDeePecker: A Deep Learning-Based System for Vulnerability DetectionZhen Li, Deqing Zou, Shouhuai Xu, Xinyu Ou 等NDSS 2018
相关 Paper
- Comparison and Evaluation of Clone Detection Techniques with Different Code RepresentationsYuekun Wang, Yuhang Ye, Yueming Wu, Weiwei Zhang 等ICSE 2023 · 被引用 15 次
- Tritor: Detecting Semantic Code Clones by Building Social Network-Based Triads ModelDeqing Zou, Siyue Feng, Yueming Wu, Wenqi Suo 等FSE 2023 · 被引用 6 次
- The Natural Geometry of Code: Hyperbolic Representation Learning for Program ReasoningWeilin ZhouICLR 2026
- Functional code clone detection with syntax and semantics fusion learningChunrong Fang, Zixi Liu, Yangyang Shi, Jeff Huang 等ISSTA 2020 · 被引用 125 次
- TreeCen: Building Tree Graph for Scalable Semantic Code Clone DetectionYutao Hu, Deqing Zou, Junru Peng, Yueming Wu 等ASE 2022 · 被引用 30 次
