PyMT5: multi-mode translation of natural language and Python code with transformers
Colin B. Clement, Dawn Drain, Jonathan Timcheck, Alexey Svyatkovskiy, Neel Sundaresan
摘要
Simultaneously modeling source code and natural language has many exciting applications in automated software development and understanding. Pursuant to achieving such technology, we introduce PYMT5, the PYTHON method text-to-text transfer transformer, which is trained to translate between all pairs of PYTHON method feature combinations: a single model that can both predict whole methods from natural language documentation strings (docstrings) and summarize code into docstrings of any common style. We present an analysis and modeling effort of a large-scale parallel corpus of 26 million PYTHON methods and 7.7 million method-docstring pairs, demonstrating that for docstring and method generation, PYMT5 outperforms similarlysized auto-regressive language models (GPT2) which were English pre-trained or randomly initialized. On the CODE-SEARCHNET test set, our best model predicts 92.1% syntactically correct method bodies, achieved a BLEU score of 8.59 for method generation and 16.3 for docstring * Corresponding author † Work done during a Microsoft internship generation (summarization), and achieved a ROUGE-L F-score of 24.8 for method generation and 36.7 for docstring generation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper19
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 被引用 1,224 次
- CodeRL: Mastering Code Generation through Pretrained Models and Deep Reinforcement LearningHung Le, Yue Wang, Akhilesh Deepak Gotmare, Silvio Savarese 等NeurIPS 2022 · 被引用 571 次
- CodeGen: An Open Large Language Model for Code with Multi-Turn Program SynthesisErik Nijkamp, Bo Pang, Hiroaki Hayashi, Lifu Tu 等ICLR 2023 · 被引用 234 次
- InCoder: A Generative Model for Code Infilling and SynthesisDaniel Fried, Armen Aghajanyan, Jessy Lin, Sida Wang 等ICLR 2023 · 被引用 140 次
- Jigsaw: Large Language Models meet Program SynthesisNaman Jain, Skanda Vaidyanath, Arun Iyer, Nagarajan Natarajan 等ICSE 2022 · 被引用 134 次
它引用的顶会 Paper1
相关 Paper
- SPT-Code: Sequence-to-Sequence Pre-Training for Learning Source Code RepresentationsChangan Niu, Chuanyi Li, Vincent Ng, Jidong Ge 等ICSE 2022 · 被引用 99 次
- Studying the Usage of Text-To-Text Transfer Transformer to Support Code-Related TasksAntonio Mastropaolo, Simone Scalabrino, Nathan Cooper, David Nader-Palacio 等ICSE 2021 · 被引用 9 次
- Long-Range Modeling of Source Code Files with eWASH: Extended Window Access by Syntax HierarchyColin B. Clement, Shuai Lu, Xiaoyu Liu, Michele Tufano 等EMNLP 2021 · 被引用 11 次
- Cell2Doc: ML Pipeline for Generating Documentation in Computational NotebooksTamal Mondal, Scott Barnett, Akash Lal, Jyothi VeduradaASE 2023 · 被引用 4 次
- Leveraging Code Generation to Improve Code Retrieval and Summarization via Dual LearningWei Ye, Rui Xie, Jinglei Zhang, Tianxiang Hu 等WWW 2020 · 被引用 83 次
