Program merge conflict resolution via neural transformers
Alexey Svyatkovskiy, Sarah Fakhoury, Negar Ghorbani, Todd Mytkowicz, Elizabeth Dinella, Christian Bird, Jinu Jang, Neel Sundaresan, Shuvendu K. Lahiri
Abstract
Collaborative software development is an integral part of the modern software development life cycle, essential to the success of largescale software projects. When multiple developers make concurrent changes around the same lines of code, a merge conflict may occur. Such conflicts stall pull requests and continuous integration pipelines for hours to several days, seriously hurting developer productivity. To address this problem, we introduce MergeBERT, a novel neural program merge framework based on token-level three-way differencing and a transformer encoder model. By exploiting the restricted nature of merge conflict resolutions, we reformulate the task of generating the resolution sequence as a classification task over a set of primitive merge patterns extracted from real-world merge commit data. Our model achieves 63-68% accuracy for merge resolution synthesis, yielding nearly a 3× performance improvement over existing semi-structured, and 2× improvement over neural program merge tools. Finally, we demonstrate that MergeBERT is sufficiently flexible to work with source code files in Java, JavaScript, Type-Script, and C# programming languages. To measure the practical use of MergeBERT, we conduct a user study to evaluate Merge-BERT suggestions with 25 developers from large OSS projects on 122 real-world conflicts they encountered. Results suggest that in practice, MergeBERT resolutions would be accepted at a higher rate than estimated by automatic metrics for precision and accuracy. Additionally, we use participant feedback to identify future avenues for improvement of MergeBERT.
• Software and its engineering → Software version control; Automatic programming.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3e2e1a9f-906c-47f4-a7ac-f30eb433e510Cited by top-tier papers4
- Contextual Predictive Mutation TestingKush Jain, Uri Alon, Alex Groce, Claire Le GouesFSE 2023 · 11 citations
- NeuDep: neural binary memory dependence analysisKexin Pei, Dongdong She, Michael Wang, Scott Geng et al.FSE 2022 · 8 citations
- Merge Conflict Resolution: Classification or Generation?Jinhao Dong, Qihao Zhu, Zeyu Sun, Yiling Lou et al.ASE 2023 · 8 citations
- Evaluation of Version Control Merge ToolsBenedikt Schesch, Ryan Featherman, Kenneth J. Yang, Ben R. Roberts et al.ASE 2024 · 1 citation
Builds on4
- Learning and Evaluating Contextual Embedding of Source CodeAditya Kanade, Petros Maniatis, Gogul Balakrishnan, Kensen ShiICML 2020 · 438 citations
- Big code != big vocabulary: open-vocabulary models for source codeRafael-Michael Karampatsis, Hlib Babii, Romain Robbes, Charles Sutton et al.ICSE 2020 · 140 citations
- BLEURT: Learning Robust Metrics for Text GenerationThibault Sellam, Dipanjan Das, Ankur P. ParikhACL 2020 · 40 citations
- Can Program Synthesis be Used to Learn Merge Conflict Resolutions? An Empirical AnalysisRangeet Pan, Vu Le, Nachiappan Nagappan, Sumit Gulwani et al.ICSE 2021 · 19 citations
Related papers
- Using pre-trained language models to resolve textual and semantic merge conflicts (experience paper)Jialu Zhang, Todd Mytkowicz, Mike Kaufman, Ruzica Piskac et al.ISSTA 2022 · 30 citations
- Planning for untangling: predicting the difficulty of merge conflictsCaius Brindescu, Iftekhar Ahmed, Rafael Leano, Anita SarmaICSE 2020 · 19 citations
- RefBERT: A Two-Stage Pre-trained Framework for Automatic Rename RefactoringHao Liu, Yanlin Wang, Zhao Wei, Yong Xu et al.ISSTA 2023 · 19 citations
- Traceability Transformed: Generating more Accurate Links with Pre-Trained BERT ModelsJinfeng Lin, Yalin Liu, Qingkai Zeng, Meng Jiang et al.ICSE 2021 · 124 citations
- An Empirical Study on Fine-Tuning Large Language Models of Code for Automated Program RepairKai Huang, Xiangxin Meng, Jian Zhang, Yang Liu et al.ASE 2023 · 91 citations
