Multilingual Code Snippets Training for Program Translation
Ming Zhu, Karthik Suresh, Chandan K. Reddy
摘要
Program translation aims to translate source code from one programming language to another. It is particularly useful in applications such as multiple-platform adaptation and legacy code migration. Traditional rule-based program translation methods usually rely on meticulous manual rule-crafting, which is costly both in terms of time and effort. Recently, neural network based methods have been developed to address this problem. However, the absence of high-quality parallel code data is one of the main bottlenecks which impedes the development of program translation models. In this paper, we introduce CoST, a new multilingual Code Snippet Translation dataset that contains parallel data from 7 commonly used programming languages. The dataset is parallel at the level of code snippets, which provides much more fine-grained alignments between different languages than the existing translation datasets. We also propose a new program translation model that leverages multilingual snippet denoising auto-encoding and Multilingual Snippet Translation (MuST) pre-training. Extensive experiments show that the multilingual snippet training is effective in improving program translation performance, especially for low-resource languages. Moreover, our training method shows good generalizability and consistently improves the translation performance of a number of baseline models. The proposed model outperforms the baselines on both snippet-level and program-level translation, and achieves state-of-the-art performance on CodeXGLUE translation task. The code, data, and appendix for this paper can be found at https://github.com/reddy-lab-code-research/MuST-CoST.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- Automating code review activities by large-scale pre-trainingZhiyu Li, Shuai Lu, Daya Guo, Nan Duan 等FSE 2022 · 被引用 195 次
- Large Language Models Meet NL2Code: A SurveyDaoguang Zan, Bei Chen, Fengji Zhang, Dianjie Lu 等ACL 2023 · 被引用 104 次
- CIRCLE: continual repair across programming languagesWei Yuan, Quanjun Zhang, Tieke He, Chunrong Fang 等ISSTA 2022 · 被引用 57 次
- Multilingual Code Co-evolution using Large Language ModelsJiyang Zhang, Pengyu Nie, Junyi Jessy Li, Milos GligoricFSE 2023 · 被引用 34 次
- Divide-and-Conquer Meets Consensus: Unleashing the Power of Functions in Code GenerationJingchang Chen, Hongxuan Tang, Zheng Chu, Qianglong Chen 等NeurIPS 2024 · 被引用 29 次
它引用的顶会 Paper4
- GraphCodeBERT: Pre-training Code Representations with Data FlowDaya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng 等ICLR 2021 · 被引用 1,644 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- Unsupervised Translation of Programming LanguagesBaptiste Rozière, Marie-Anne Lachaux, Lowik Chanussot, Guillaume LampleNeurIPS 2020 · 被引用 606 次
- DOBF: A Deobfuscation Pre-Training Objective for Programming LanguagesMarie-Anne Lachaux, Baptiste Rozière, Marc Szafraniec, Guillaume LampleNeurIPS 2021 · 被引用 174 次
相关 Paper
- Program Translation via Code DistillationYufan Huang, Mengnan Qi, Yongqiang Yao, Maoquan Wang 等EMNLP 2023 · 被引用 7 次
- AdaCCD: Adaptive Semantic Contrasts Discovery Based Cross Lingual Adaptation for Code Clone DetectionYangkai Du, Tengfei Ma, Lingfei Wu, Xuhong Zhang 等AAAI 2024 · 被引用 9 次
- Code–Test Co-translation: Towards Practical and Effective Program Migration in the WildXitao Li, Xiaofei Xie, Jiang Wu, Ting Liu 等OOPSLA 2026 · 被引用 1 次
- The Struggles of LLMs in Cross-Lingual Code Clone DetectionMicheline Bénédicte Moumoula, Abdoul Kader Kaboré, Jacques Klein, Tegawendé F. BissyandéFSE 2025 · 被引用 4 次
- Syntax and Domain Aware Model for Unsupervised Program TranslationFang Liu, Jia Li, Li ZhangICSE 2023 · 被引用 25 次
