Two Birds with One Stone: Boosting Code Generation and Code Search via a Generative Adversarial Network
Shangwen Wang, Bo Lin, Zhensu Sun, Ming Wen, Yepang Liu, Yan Lei, Xiaoguang Mao
摘要
Automatically transforming developers' natural language descriptions into source code has been a longstanding goal in software engineering research. Two types of approaches have been proposed in the literature to achieve this: code generation, which involves generating a new code snippet, and code search, which involves reusing existing code. However, despite existing efforts, the effectiveness of the state-of-the-art techniques remains limited. To seek for further advancement, our insight is that code generation and code search can help overcome the limitation of each other: the code generator can benefit from feedback on the quality of its generated code, which can be provided by the code searcher, while the code searcher can benefit from the additional training data augmented by the code generator to better understand code semantics. Drawing on this insight, we propose a novel approach that combines code generation and code search techniques using a generative adversarial network (GAN), enabling mutual improvement through the adversarial training. Specifically, we treat code generation and code search as the generator and discriminator in the GAN framework, respectively, and incorporate several customized designs for our tasks. We evaluate our approach in eight different settings, and consistently observe significant performance improvements for both code generation and code search. For instance, when using NatGen, a state-of-the-art code generator, as the generator and GraphCodeBERT, a state-of-the-art code searcher, as the discriminator, we achieve a 32% increase in CodeBLEU score for code generation, and a 12% increase in mean reciprocal rank for code search on a large-scale Python dataset, compared to their original performances.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Instruct or Interact? Exploring and Eliciting LLMs' Capability in Code Snippet Adaptation Through Prompt EngineeringTanghaoran Zhang, Yue Yu, Xinjun Mao, Shangwen Wang 等ICSE 2025 · 被引用 3 次
- 10 Years Later: Revisiting How Developers Search for CodeKathryn T. Stolee, Tobias Welp, Caitlin Sadowski, Sebastian G. ElbaumFSE 2025 · 被引用 3 次
- Model Editing for LLMs4Code: How Far are we?Xiaopeng Li, Shangwen Wang, Shasha Li, Jun Ma 等ICSE 2025 · 被引用 2 次
- LLM-based Agents for Automated Bug Fixing: How Far Are We?Xiangxin Meng, Zexiong Ma, Pengfei Gao, Chao PengICSE 2026
它引用的顶会 Paper20
- GraphCodeBERT: Pre-training Code Representations with Data FlowDaya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng 等ICLR 2021 · 被引用 1,644 次
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 被引用 1,224 次
- Asleep at the Keyboard? Assessing the Security of GitHub Copilot's Code ContributionsHammond Pearce, Baleegh Ahmad, Benjamin Tan, Brendan Dolan-Gavitt 等S&P 2022 · 被引用 725 次
- Grounded Copilot: How Programmers Interact with Code-Generating ModelsShraddha Barke, Michael B. James, Nadia PolikarpovaOOPSLA 2023 · 被引用 408 次
- Automated Repair of Programs from Large Language ModelsZhiyu Fan, Xiang Gao, Martin Mirchev, Abhik Roychoudhury 等ICSE 2023 · 被引用 213 次
相关 Paper
- Natural Language to Code: How Far Are We?Shangwen Wang, Mingyang Geng, Bo Lin, Zhensu Sun 等FSE 2023 · 被引用 24 次
- Leveraging Code Generation to Improve Code Retrieval and Summarization via Dual LearningWei Ye, Rui Xie, Jinglei Zhang, Tianxiang Hu 等WWW 2020 · 被引用 83 次
- CodeGen4Libs: A Two-Stage Approach for Library-Oriented Code GenerationMingwei Liu, Tianyong Yang, Yiling Lou, Xueying Du 等ASE 2023 · 被引用 31 次
- SkCoder: A Sketch-based Approach for Automatic Code GenerationJia Li, Yongmin Li, Ge Li, Zhi Jin 等ICSE 2023 · 被引用 50 次
- PyMT5: multi-mode translation of natural language and Python code with transformersColin B. Clement, Dawn Drain, Jonathan Timcheck, Alexey Svyatkovskiy 等EMNLP 2020 · 被引用 24 次
