Program Translation via Code Distillation
Yufan Huang, Mengnan Qi, Yongqiang Yao, Maoquan Wang, Bin Gu, Colin B. Clement, Neel Sundaresan
摘要
Software version migration and program translation are an important and costly part of the lifecycle of large codebases. Traditional machine translation relies on parallel corpora for supervised translation, which is not feasible for program translation due to a dearth of aligned data. Recent unsupervised neural machine translation techniques have overcome data limitations by included techniques such as back translation and low level compiler intermediate representations (IR). These methods face significant challenges due to the noise in code snippet alignment and the diversity of IRs respectively. In this paper we propose a novel model called Code Distillation (CoDist) whereby we capture the semantic and structural equivalence of code in a language agnostic intermediate representation. Distilled code serves as a translation pivot for any programming language, leading by construction to parallel corpora which scale to all available source code by simply applying the distillation compiler. We demonstrate that our approach achieves state-of-the-art performance on CodeXGLUE and TransCoder GeeksForGeeks translation benchmarks, with an average absolute increase of 12.7% on the TransCoder GeeksforGeeks translation benchmark compare to TransCoder-ST.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Learning to Focus: Causal Attention Distillation via Gradient-Guided Token PruningYiju Guo, Wenkai Yang, Zexu Sun, Ning Ding 等NeurIPS 2025 · 被引用 14 次
- AI Coders Are among Us: Rethinking Programming Language Grammar towards Efficient Code GenerationZhensu Sun, Xiaoning Du, Zhou Yang, Li Li 等ISSTA 2024 · 被引用 8 次
- On-Policy Optimization with Group Equivalent Preference for Multi-Programming Language UnderstandingHaoyuan Wu, Rui Ming, Jilong Gao, Hangyu Zhao 等NeurIPS 2025 · 被引用 2 次
- Semi-Supervised Code Translation Overcoming the Scarcity of Parallel Code DataMing Zhu, Mohimenul Karim, Ismini Lourentzou, Daphne YaoASE 2024 · 被引用 2 次
- Token Sugar: Making Source Code Sweeter for LLMs through Token-Efficient ShorthandZhensu Sun, Chengran Yang, Xiaoning Du, Zhou Yang 等ASE 2025 · 被引用 2 次
它引用的顶会 Paper8
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- GraphCodeBERT: Pre-training Code Representations with Data FlowDaya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng 等ICLR 2021 · 被引用 1,644 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- Unsupervised Translation of Programming LanguagesBaptiste Rozière, Marie-Anne Lachaux, Lowik Chanussot, Guillaume LampleNeurIPS 2020 · 被引用 606 次
- Leveraging Automated Unit Tests for Unsupervised Code TranslationBaptiste Rozière, Jie Zhang, François Charton, Mark Harman 等ICLR 2022 · 被引用 161 次
相关 Paper
- Code Translation with Compiler RepresentationsMarc Szafraniec, Baptiste Rozière, Hugh Leather, Patrick Labatut 等ICLR 2023 · 被引用 18 次
- Multilingual Code Snippets Training for Program TranslationMing Zhu, Karthik Suresh, Chandan K. ReddyAAAI 2022 · 被引用 72 次
- Syntax and Domain Aware Model for Unsupervised Program TranslationFang Liu, Jia Li, Li ZhangICSE 2023 · 被引用 25 次
- Multilingual Code Co-evolution using Large Language ModelsJiyang Zhang, Pengyu Nie, Junyi Jessy Li, Milos GligoricFSE 2023 · 被引用 34 次
- INTERTRANS: Leveraging Transitive Intermediate Translations to Enhance LLM-Based Code TranslationMarcos Macedo, Yuan Tian, Pengyu Nie, Filipe Roseiro Côgo 等ICSE 2025 · 被引用 7 次
