Learning Structural Edits via Incremental Tree Transformations
Ziyu Yao, Frank F. Xu, Pengcheng Yin, Huan Sun, Graham Neubig
摘要
While most neural generative models generate outputs in a single pass, the human creative process is usually one of iterative building and refinement. Recent work has proposed models of editing processes, but these mostly focus on editing sequential data and/or only model a single editing pass. In this paper, we present a generic model for incremental editing of structured data (i.e. ''structural edits''). Particularly, we focus on tree-structured data, taking abstract syntax trees of computer programs as our canonical example. Our editor learns to iteratively generate tree edits (e.g. deleting or adding a subtree) and applies them to the partially edited data, thereby the entire editing process can be formulated as consecutive, incremental tree transformations. To show the unique benefits of modeling tree edits directly, we further propose a novel edit encoder for learning to represent edits, as well as an imitation learning method that allows the editor to be more robust. We evaluate our proposed editor on two source code edit datasets, where results show that, with the proposed edit encoder, our editor significantly improves accuracy over previous approaches that generate the edited program directly in one pass. Finally, we demonstrate that training our editor to imitate experts and correct its mistakes dynamically can further improve its performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Aligning LLM Agents by Learning Latent Preference from User EditsGe Gao, Alexey Taymanov, Eduardo Salinas, Paul Mineiro 等NeurIPS 2024 · 被引用 102 次
- CoditT5: Pretraining for Source Code and Natural Language EditingJiyang Zhang, Sheena Panthaplackel, Pengyu Nie, Junyi Jessy Li 等ASE 2022 · 被引用 81 次
- A Deep Dive into Large Language Models for Automated Bug Localization and RepairSoneya Binta Hossain, Nan Jiang, Qiang Zhou, Xiaopeng Li 等FSE 2024 · 被引用 60 次
- KNOD: Domain Knowledge Distilled Tree Decoder for Automated Program RepairNan Jiang, Thibaud Lutellier, Yiling Lou, Lin Tan 等ICSE 2023 · 被引用 55 次
- Multilingual Code Co-evolution using Large Language ModelsJiyang Zhang, Pengyu Nie, Junyi Jessy Li, Milos GligoricFSE 2023 · 被引用 34 次
它引用的顶会 Paper11
- Global Relational Models of Source CodeVincent J. Hellendoorn, Charles Sutton, Rishabh Singh, Petros Maniatis 等ICLR 2020 · 被引用 252 次
- Hoppity: Learning Graph Transformations to Detect and Fix Bugs in ProgramsElizabeth Dinella, Hanjun Dai, Ziyang Li, Mayur Naik 等ICLR 2020 · 被引用 212 次
- Graph-based, Self-Supervised Program Repair from Diagnostic FeedbackMichihiro Yasunaga, Percy LiangICML 2020 · 被引用 198 次
- CC2Vec: distributed representations of code changesThong Hoang, Hong Jin Kang, David Lo, Julia LawallICSE 2020 · 被引用 169 次
- A structural model for contextual code changesShaked Brody, Uri Alon, Eran YahavOOPSLA 2020 · 被引用 36 次
相关 Paper
- An Imitation Learning Curriculum for Text Editing with Non-Autoregressive ModelsSweta Agrawal, Marine CarpuatACL 2022
- Diffusion On Syntax Trees For Program SynthesisShreyas Kapur, Erik Jenner, Stuart RussellICLR 2025
- NextCoder: Robust Adaptation of Code LMs to Diverse Code EditsTushar Aggarwal, Swayam Singh, Abhijeet Awasthi, Aditya Kanade 等ICML 2025
- Towards Robust Sequential Decomposition for Complex Image EditingZilai Zeng, Mingdeng Cao, Zijie Li, Xiaochen Lian 等CVPR 2026
- Interactive Text GenerationFelix Faltings, Michel Galley, Kianté Brantley, Baolin Peng 等EMNLP 2023 · 被引用 7 次
