Multilingual Code Co-evolution using Large Language Models
Jiyang Zhang, Pengyu Nie, Junyi Jessy Li, Milos Gligoric
摘要
Many software projects implement APIs and algorithms in multiple programming languages. Maintaining such projects is tiresome, as developers have to ensure that any change (e.g., a bug fix or a new feature) is being propagated, timely and without errors, to implementations in other programming languages. In the world of ever-changing software, using rule-based translation tools (i.e., transpilers) or machine learning models for translating code from one language to another provides limited value. Translating each time the entire codebase from one language to another is not the way developers work. In this paper, we target a novel task: translating code changes from one programming language to another using large language models (LLMs). We design and implement the first LLM, dubbed Codeditor, to tackle this task. Codeditor explicitly models code changes as edit sequences and learns to correlate changes across programming languages. To evaluate Codeditor, we collect a corpus of 6,613 aligned code changes from 8 pairs of open-source software projects implementing similar functionalities in two programming languages (Java and C#). Results show that Codeditor outperforms the state-of-the-art approaches by a large margin on all commonly used automatic metrics. Our work also reveals that Codeditor is complementary to the existing generation-based models, and their combination ensures even greater performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Scalable, Validated Code Translation of Entire Projects using Large Language ModelsHanliang Zhang, Cristina David, Meng Wang, Brandon Paulsen 等PLDI 2025 · 被引用 15 次
- An interpretable error correction method for enhancing code-to-code translationMin Xue, Artur Andrzejak, Marla LeutherICLR 2024 · 被引用 10 次
- If At First You Don't Succeed, Try, Try, Again...? Insights and LLM-informed Tooling for Detecting Retry Bugs in Software SystemsBogdan Alexandru Stoica, Utsav Sethi, Yiming Su, Cyrus Zhou 等SOSP 2024 · 被引用 4 次
- Promise and Peril of Collaborative Code Generation Models: Balancing Effectiveness and MemorizationZhi Chen, Lingxiao JiangASE 2024 · 被引用 4 次
- exLong: Generating Exceptional Behavior Tests with Large Language ModelsJiyang Zhang, Yu Liu, Pengyu Nie, Junyi Jessy Li 等ICSE 2025 · 被引用 2 次
它引用的顶会 Paper21
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 被引用 1,224 次
- Unsupervised Translation of Programming LanguagesBaptiste Rozière, Marie-Anne Lachaux, Lowik Chanussot, Guillaume LampleNeurIPS 2020 · 被引用 606 次
- CodeT5+: Open Code Large Language Models for Code Understanding and GenerationYue Wang, Hung Le, Akhilesh Gotmare, Nghi D. Q. Bui 等EMNLP 2023 · 被引用 339 次
相关 Paper
- INTERTRANS: Leveraging Transitive Intermediate Translations to Enhance LLM-Based Code TranslationMarcos Macedo, Yuan Tian, Pengyu Nie, Filipe Roseiro Côgo 等ICSE 2025 · 被引用 7 次
- Polyglot: An Extensible Framework to Benchmark Code Translation with LLMsMarco Vieira, Priyam Ashish Shah, Bhavain Shah, Rrezarta KrasniqiASE 2025
- Learning to Update Natural Language Comments Based on Code ChangesSheena Panthaplackel, Pengyu Nie, Milos Gligoric, Junyi Jessy Li 等ACL 2020 · 被引用 1 次
- ExeCoder: Empowering Large Language Models with Executability Representation for Code TranslationMinghua He, Yue Chen, Fangkai Yang, Pu Zhao 等EMNLP 2025 · 被引用 1 次
- Lost in Translation: A Study of Bugs Introduced by Large Language Models while Translating CodeRangeet Pan, Ali Reza Ibrahimzada, Rahul Krishna, Divya Sankar 等ICSE 2024 · 被引用 96 次
