CodeTrek: Flexible Modeling of Code using an Extensible Relational Representation
Pardis Pashakhanloo, Aaditya Naik, Yuepeng Wang, Hanjun Dai, Petros Maniatis, Mayur Naik
摘要
Designing a suitable representation for code-reasoning tasks is challenging in aspects such as the kinds of program information to model, how to combine them, and how much context to consider. We propose CodeTrek, a deep learning approach that addresses these challenges by representing codebases as databases that conform to rich relational schemas. The relational representation not only allows CodeTrek to uniformly represent diverse kinds of program information, but also to leverage program-analysis queries to derive new semantic relations, which can be readily incorporated without further architectural engineering. CodeTrek embeds this relational representation using a set of walks that can traverse different relations in an unconstrained fashion, and incorporates all relevant attributes along the way. We evaluate CodeTrek on four diverse and challenging Python tasks: variable misuse, exception prediction, unused definition, and variable shadowing. CodeTrek achieves an accuracy of 91%, 63%, 98%, and 94% on these tasks respectively, and outperforms state-of-the-art neural models by 2-19% points.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- CodePlan: Repository-Level Coding using LLMs and PlanningRamakrishna Bairi, Atharv Sonwane, Aditya Kanade, Vageesh D. C. 等FSE 2024 · 被引用 67 次
- On Grounded Planning for Embodied Tasks with Language ModelsBill Yuchen Lin, Chengsong Huang, Qian Liu, Wenda Gu 等AAAI 2023 · 被引用 52 次
- Statically Contextualizing Large Language Models with Typed HolesAndrew Blinn, Xiang Li, June Hyung Kim, Cyrus OmarOOPSLA 2024 · 被引用 10 次
它引用的顶会 Paper12
- Learning and Evaluating Contextual Embedding of Source CodeAditya Kanade, Petros Maniatis, Gogul Balakrishnan, Kensen ShiICML 2020 · 被引用 438 次
- Global Relational Models of Source CodeVincent J. Hellendoorn, Charles Sutton, Rishabh Singh, Petros Maniatis 等ICLR 2020 · 被引用 252 次
- Hoppity: Learning Graph Transformations to Detect and Fix Bugs in ProgramsElizabeth Dinella, Hanjun Dai, Ziyang Li, Mayur Naik 等ICLR 2020 · 被引用 212 次
- Code Prediction by Feeding Trees to TransformersSeohyun Kim, Jinman Zhao, Yuchi Tian, Satish ChandraICSE 2021 · 被引用 179 次
- Multi-task Learning based Pre-trained Language Model for Code CompletionFang Liu, Ge Li, Yunfei Zhao, Zhi JinASE 2020 · 被引用 162 次
相关 Paper
- TreeCaps: Tree-Based Capsule Networks for Source Code ProcessingNghi D. Q. Bui, Yijun Yu, Lingxiao JiangAAAI 2021 · 被引用 44 次
- Fault Localization with Code Coverage Representation LearningYi Li, Shaohua Wang, Tien N. NguyenICSE 2021 · 被引用 120 次
- Blended, precise semantic program embeddingsKe Wang, Zhendong SuPLDI 2020 · 被引用 53 次
- Keeping Pace with Ever-Increasing Data: Towards Continual Learning of Code Intelligence ModelsShuzheng Gao, Hongyu Zhang, Cuiyun Gao, Chaozheng WangICSE 2023 · 被引用 14 次
- Static Inference Meets Deep learning: A Hybrid Type Inference Approach for PythonYun Peng, Cuiyun Gao, Zongjie Li, Bowei Gao 等ICSE 2022 · 被引用 48 次
