Learning Structural Edits via Incremental Tree Transformations
Ziyu Yao, Frank F. Xu, Pengcheng Yin, Huan Sun, Graham Neubig
Abstract
While most neural generative models generate outputs in a single pass, the human creative process is usually one of iterative building and refinement. Recent work has proposed models of editing processes, but these mostly focus on editing sequential data and/or only model a single editing pass. In this paper, we present a generic model for incremental editing of structured data (i.e. ''structural edits''). Particularly, we focus on tree-structured data, taking abstract syntax trees of computer programs as our canonical example. Our editor learns to iteratively generate tree edits (e.g. deleting or adding a subtree) and applies them to the partially edited data, thereby the entire editing process can be formulated as consecutive, incremental tree transformations. To show the unique benefits of modeling tree edits directly, we further propose a novel edit encoder for learning to represent edits, as well as an imitation learning method that allows the editor to be more robust. We evaluate our proposed editor on two source code edit datasets, where results show that, with the proposed edit encoder, our editor significantly improves accuracy over previous approaches that generate the edited program directly in one pass. Finally, we demonstrate that training our editor to imitate experts and correct its mistakes dynamically can further improve its performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 01fe4680-9ef2-4dac-9e11-4f9ae116f31dCited by top-tier papers11
- Aligning LLM Agents by Learning Latent Preference from User EditsGe Gao, Alexey Taymanov, Eduardo Salinas, Paul Mineiro et al.NeurIPS 2024 · 102 citations
- CoditT5: Pretraining for Source Code and Natural Language EditingJiyang Zhang, Sheena Panthaplackel, Pengyu Nie, Junyi Jessy Li et al.ASE 2022 · 81 citations
- A Deep Dive into Large Language Models for Automated Bug Localization and RepairSoneya Binta Hossain, Nan Jiang, Qiang Zhou, Xiaopeng Li et al.FSE 2024 · 60 citations
- KNOD: Domain Knowledge Distilled Tree Decoder for Automated Program RepairNan Jiang, Thibaud Lutellier, Yiling Lou, Lin Tan et al.ICSE 2023 · 55 citations
- Multilingual Code Co-evolution using Large Language ModelsJiyang Zhang, Pengyu Nie, Junyi Jessy Li, Milos GligoricFSE 2023 · 34 citations
Builds on11
- Global Relational Models of Source CodeVincent J. Hellendoorn, Charles Sutton, Rishabh Singh, Petros Maniatis et al.ICLR 2020 · 252 citations
- Hoppity: Learning Graph Transformations to Detect and Fix Bugs in ProgramsElizabeth Dinella, Hanjun Dai, Ziyang Li, Mayur Naik et al.ICLR 2020 · 212 citations
- Graph-based, Self-Supervised Program Repair from Diagnostic FeedbackMichihiro Yasunaga, Percy LiangICML 2020 · 198 citations
- CC2Vec: distributed representations of code changesThong Hoang, Hong Jin Kang, David Lo, Julia LawallICSE 2020 · 169 citations
- A structural model for contextual code changesShaked Brody, Uri Alon, Eran YahavOOPSLA 2020 · 36 citations
Related papers
- An Imitation Learning Curriculum for Text Editing with Non-Autoregressive ModelsSweta Agrawal, Marine CarpuatACL 2022
- Diffusion On Syntax Trees For Program SynthesisShreyas Kapur, Erik Jenner, Stuart RussellICLR 2025
- NextCoder: Robust Adaptation of Code LMs to Diverse Code EditsTushar Aggarwal, Swayam Singh, Abhijeet Awasthi, Aditya Kanade et al.ICML 2025
- Towards Robust Sequential Decomposition for Complex Image EditingZilai Zeng, Mingdeng Cao, Zijie Li, Xiaochen Lian et al.CVPR 2026
- Interactive Text GenerationFelix Faltings, Michel Galley, Kianté Brantley, Baolin Peng et al.EMNLP 2023 · 7 citations
