PinNet: Pinpoint Instructive Information for Retrieval Augmented Code-to-Text Generation
Han Fu, Jian Tan, Pinhan Zhang, Feifei Li, Jianling Sun
Abstract
Automatically generating high quality code descriptions greatly improves the readability and maintainability of the codebase. Recently, retrieval augmented code-to-text approaches have proven to be an effective solution, and have achieved the state-of-the-art results on various benchmarks. It brings out the potential to leverage large unlabeled code descriptions to further improve the generation quality. In spite of the promising performance, retrieval-augmented models however suffer from being deluded by inconducive retrieved references, due to irrelevant or even misleading information contained therein. To this end, we design PinNet, a new framework for code-to-text generation. PinNet relies on a discriminator to measure how well the retrievals match the semantics of the input code. Remarkably, the hidden representation of the reference from the last layer of the discriminator can be leveraged to significantly improve the codeto-text generation through modifying the attention weights. It essentially pays high attention to valuable information and eliminates misleading part. To effectively execute this idea, we also propose a novel contrastive learning method to quantify the semantical similarities between unlabeled references. Using extensive experiments on code summarization and SQL-to-text generation, we demonstrate that the proposed method can significantly outperform all of the baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dc87a62d-e487-488f-ba28-eed9436b93ddBuilds on12
- Understanding Contrastive Representation Learning through Alignment and Uniformity on the HypersphereTongzhou Wang, Phillip IsolaICML 2020 · 2,360 citations
- GraphCodeBERT: Pre-training Code Representations with Data FlowDaya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng et al.ICLR 2021 · 1,644 citations
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 1,224 citations
- Retrieval-based neural source code summarizationJian Zhang, Xu Wang, Hongyu Zhang, Hailong Sun et al.ICSE 2020 · 242 citations
- CodeGen: An Open Large Language Model for Code with Multi-Turn Program SynthesisErik Nijkamp, Bo Pang, Hiroaki Hayashi, Lifu Tu et al.ICLR 2023 · 234 citations
Related papers
- Self-Supervised Contrastive Learning for Code Retrieval and Summarization via Semantic-Preserving TransformationsNghi D. Q. Bui, Yijun Yu, Lingxiao JiangSIGIR 2021 · 98 citations
- CodeRetriever: A Large Scale Contrastive Pre-Training Method for Code SearchXiaonan Li, Yeyun Gong, Yelong Shen, Xipeng Qiu et al.EMNLP 2022 · 25 citations
- CoNT: Contrastive Neural Text GenerationChenxin An, Jiangtao Feng, Kai Lv, Lingpeng Kong et al.NeurIPS 2022 · 37 citations
- Evaluating Code Summarization Techniques: A New Metric and an Empirical CharacterizationAntonio Mastropaolo, Matteo Ciniselli, Massimiliano Di Penta, Gabriele BavotaICSE 2024 · 28 citations
- Retrieve and Refine: Exemplar-based Neural Comment GenerationBolin Wei, Yongmin Li, Ge Li, Xin Xia et al.ASE 2020 · 68 citations
