Neural Program Generation Modulo Static Analysis
Rohan Mukherjee, Yeming Wen, Dipak Chaudhari, Thomas W. Reps, Swarat Chaudhuri, Christopher M. Jermaine
摘要
State-of-the-art neural models of source code tend to be evaluated on the generation of individual expressions and lines of code, and commonly fail on long-horizon tasks such as the generation of entire method bodies. We propose to address this deficiency using weak supervision from a static program analyzer. Our neurosymbolic method allows a deep generative model to symbolically compute, using calls to a static-analysis tool, long-distance semantic relationships in the code that it has already generated. During training, the model observes these relationships and learns to generate programs conditioned on them. We apply our approach to the problem of generating entire Java methods given the remainder of the class that contains the method. Our experiments show that the approach substantially outperforms state-of-the-art transformers and a model that explicitly tries to learn program semantics on this task, both in terms of producing programs free of basic semantic errors and in terms of syntactically matching the ground truth.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Fault-Aware Neural Code RankersJeevana Priya Inala, Chenglong Wang, Mei Yang, Andrés Codas 等NeurIPS 2022 · 被引用 61 次
- Monitor-Guided Decoding of Code LMs with Static Analysis of Repository ContextLakshya A. Agrawal, Aditya Kanade, Navin Goyal, Shuvendu K. Lahiri 等NeurIPS 2023 · 被引用 55 次
- On Distribution Shift in Learning-based Bug DetectorsJingxuan He, Luca Beurer-Kellner, Martin T. VechevICML 2022 · 被引用 20 次
- Rethinking Repetition Problems of LLMs in Code GenerationYihong Dong, Yuchen Liu, Xue Jiang, Bin Gu 等ACL 2025
- XFL: Naming Functions in Binaries with Extreme Multi-label LearningJames Patrick-Evans, Moritz Dannehl, Johannes KinderS&P 2023
它引用的顶会 Paper3
- Learning Differentiable Programs with Admissible Neural HeuristicsAmeesh Shah, Eric Zhan, Jennifer J. Sun, Abhinav Verma 等NeurIPS 2020 · 被引用 56 次
- Learning to Represent Programs with Property SignaturesAugustus Odena, Charles SuttonICLR 2020 · 被引用 34 次
- Program Synthesis Using Deduction-Guided Reinforcement LearningYanju Chen, Chenglong Wang, Osbert Bastani, Isil Dillig 等CAV 2020 · 被引用 30 次
相关 Paper
- Web question answering with neurosymbolic program synthesisQiaochu Chen, Aaron Lamoreaux, Xinyu Wang, Greg Durrett 等PLDI 2021 · 被引用 25 次
- Guess & Sketch: Language Model Guided TranspilationCeline Lee, Abdulrahman Mahmoud, Michal Kurek, Simone Campanoni 等ICLR 2024 · 被引用 8 次
- Representing Partial Programs with Blended Abstract SemanticsMaxwell I. Nye, Yewen Pu, Matthew Bowers, Jacob Andreas 等ICLR 2021 · 被引用 23 次
- Blended, precise semantic program embeddingsKe Wang, Zhendong SuPLDI 2020 · 被引用 53 次
- Towards Reliable Code-as-Policies: A Neuro-Symbolic Framework for Embodied Task PlanningSanghyun Ahn, Wonje Choi, Junyong Lee, Jinwoo Park 等NeurIPS 2025 · 被引用 14 次
