Multilingual Code Snippets Training for Program Translation
Ming Zhu, Karthik Suresh, Chandan K. Reddy
Abstract
Program translation aims to translate source code from one programming language to another. It is particularly useful in applications such as multiple-platform adaptation and legacy code migration. Traditional rule-based program translation methods usually rely on meticulous manual rule-crafting, which is costly both in terms of time and effort. Recently, neural network based methods have been developed to address this problem. However, the absence of high-quality parallel code data is one of the main bottlenecks which impedes the development of program translation models. In this paper, we introduce CoST, a new multilingual Code Snippet Translation dataset that contains parallel data from 7 commonly used programming languages. The dataset is parallel at the level of code snippets, which provides much more fine-grained alignments between different languages than the existing translation datasets. We also propose a new program translation model that leverages multilingual snippet denoising auto-encoding and Multilingual Snippet Translation (MuST) pre-training. Extensive experiments show that the multilingual snippet training is effective in improving program translation performance, especially for low-resource languages. Moreover, our training method shows good generalizability and consistently improves the translation performance of a number of baseline models. The proposed model outperforms the baselines on both snippet-level and program-level translation, and achieves state-of-the-art performance on CodeXGLUE translation task. The code, data, and appendix for this paper can be found at https://github.com/reddy-lab-code-research/MuST-CoST.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8d473930-1127-4095-98f3-e111076b22f5Cited by top-tier papers17
- Automating code review activities by large-scale pre-trainingZhiyu Li, Shuai Lu, Daya Guo, Nan Duan et al.FSE 2022 · 195 citations
- Large Language Models Meet NL2Code: A SurveyDaoguang Zan, Bei Chen, Fengji Zhang, Dianjie Lu et al.ACL 2023 · 104 citations
- CIRCLE: continual repair across programming languagesWei Yuan, Quanjun Zhang, Tieke He, Chunrong Fang et al.ISSTA 2022 · 57 citations
- Multilingual Code Co-evolution using Large Language ModelsJiyang Zhang, Pengyu Nie, Junyi Jessy Li, Milos GligoricFSE 2023 · 34 citations
- Divide-and-Conquer Meets Consensus: Unleashing the Power of Functions in Code GenerationJingchang Chen, Hongxuan Tang, Zheng Chu, Qianglong Chen et al.NeurIPS 2024 · 29 citations
Builds on4
- GraphCodeBERT: Pre-training Code Representations with Data FlowDaya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng et al.ICLR 2021 · 1,644 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- Unsupervised Translation of Programming LanguagesBaptiste Rozière, Marie-Anne Lachaux, Lowik Chanussot, Guillaume LampleNeurIPS 2020 · 606 citations
- DOBF: A Deobfuscation Pre-Training Objective for Programming LanguagesMarie-Anne Lachaux, Baptiste Rozière, Marc Szafraniec, Guillaume LampleNeurIPS 2021 · 174 citations
Related papers
- Program Translation via Code DistillationYufan Huang, Mengnan Qi, Yongqiang Yao, Maoquan Wang et al.EMNLP 2023 · 7 citations
- AdaCCD: Adaptive Semantic Contrasts Discovery Based Cross Lingual Adaptation for Code Clone DetectionYangkai Du, Tengfei Ma, Lingfei Wu, Xuhong Zhang et al.AAAI 2024 · 9 citations
- Code–Test Co-translation: Towards Practical and Effective Program Migration in the WildXitao Li, Xiaofei Xie, Jiang Wu, Ting Liu et al.OOPSLA 2026 · 1 citation
- The Struggles of LLMs in Cross-Lingual Code Clone DetectionMicheline Bénédicte Moumoula, Abdoul Kader Kaboré, Jacques Klein, Tegawendé F. BissyandéFSE 2025 · 4 citations
- Syntax and Domain Aware Model for Unsupervised Program TranslationFang Liu, Jia Li, Li ZhangICSE 2023 · 25 citations
