Lune

ASE2022Top-tier venue

CoditT5: Pretraining for Source Code and Natural Language Editing

Jiyang Zhang, Sheena Panthaplackel, Pengyu Nie, Junyi Jessy Li, Milos Gligoric

2022Year
81Citations
28Top-tier citations

Abstract

Pretrained language models have been shown to be effective in many software-related generation tasks; however, they are not wellsuited for editing tasks as they are not designed to reason about edits. To address this, we propose a novel pretraining objective which explicitly models edits and use it to build CoditT5, a large language model for software-related editing tasks that is pretrained on large amounts of source code and natural language comments. We fine-tune it on various downstream editing tasks, including comment updating, bug fixing, and automated code review. By outperforming standard generation-based models, we demonstrate the generalizability of our approach and its suitability for editing tasks. We also show how a standard generation model and our editbased model can complement one another through simple reranking strategies, with which we achieve state-of-the-art performance for the three downstream editing tasks. CCS CONCEPTS • Computing methodologies → Machine learning; • Software and its engineering → Software evolution.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext df818fec-e067-460c-95e0-614878135b98

Cited by top-tier papers28

Ask how each one uses it

Builds on19

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines