Semi-Supervised Code Translation Overcoming the Scarcity of Parallel Code Data
Ming Zhu, Mohimenul Karim, Ismini Lourentzou, Daphne Yao
Abstract
Neural code translation is the task of converting source code from one programming language to another. One of the main challenges is the scarcity of parallel code data, which hinders the ability of translation models to learn accurate cross-language alignments. In this paper, we introduce MIRACLE, a semi-supervised approach that improves code translation through synthesizing high-quality parallel code data and curriculum learning on code data with ascending alignment levels. MIRACLE leverages static analysis and compilation to generate synthetic parallel code datasets with enhanced quality and alignment to address the challenge of data scarcity. We evaluate the proposed method along with strong baselines including instruction-tuned Large Language Models (LLMs) for code. Our analysis reveals that LLMs pre-trained on open-source code data, regardless of their size, suffer from the "shallow translation" problem. This issue arises when translated code copies keywords, statements, and even code blocks from the source language, leading to compilation and runtime errors. Extensive experiments demonstrate that our method significantly mitigates this issue, enhancing code translation performance across multiple models in C++, Java, Python, and C. Remarkably, MIRACLE outperforms code LLMs that are ten times larger in size. MIRACLE also achieves up to a 43% improvement in C code translation with fewer than 150 annotated examples.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6e3468da-c2e4-4aee-aff1-ee8f50954f69Cited by top-tier papers4
- TRACE: Evaluating Execution Efficiency of LLM-Based Code TranslationZhihao Gong, Zeyu Sun, Dong Huang, Qingyuan Liang et al.ACL 2026 · 5 citations
- QiMeng-MuPa: Mutual-Supervised Learning for Sequential-to-Parallel Code TranslationChangxin Ke, Rui Zhang, Shuo Wang, Li Ding et al.NeurIPS 2025 · 3 citations
- ExeCoder: Empowering Large Language Models with Executability Representation for Code TranslationMinghua He, Yue Chen, Fangkai Yang, Pu Zhao et al.EMNLP 2025 · 1 citation
- Bootstrapping Code Translation with Weighted Multilanguage ExplorationYuhan Wu, Huan Zhang, Wei Cheng, Chen Shen et al.ACL 2026
Builds on14
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- GraphCodeBERT: Pre-training Code Representations with Data FlowDaya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng et al.ICLR 2021 · 1,644 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 1,224 citations
- Unsupervised Translation of Programming LanguagesBaptiste Rozière, Marie-Anne Lachaux, Lowik Chanussot, Guillaume LampleNeurIPS 2020 · 606 citations
Related papers
- Lost in Translation: A Study of Bugs Introduced by Large Language Models while Translating CodeRangeet Pan, Ali Reza Ibrahimzada, Rahul Krishna, Divya Sankar et al.ICSE 2024 · 96 citations
- Knowledge Transfer from High-Resource to Low-Resource Programming Languages for Code LLMsFederico Cassano, John Gouwar, Francesca Lucchetti, Claire Schlesinger et al.OOPSLA 2024 · 33 citations
- Program Translation via Code DistillationYufan Huang, Mengnan Qi, Yongqiang Yao, Maoquan Wang et al.EMNLP 2023 · 7 citations
- Code Translation with Compiler RepresentationsMarc Szafraniec, Baptiste Rozière, Hugh Leather, Patrick Labatut et al.ICLR 2023 · 18 citations
- INTERTRANS: Leveraging Transitive Intermediate Translations to Enhance LLM-Based Code TranslationMarcos Macedo, Yuan Tian, Pengyu Nie, Filipe Roseiro Côgo et al.ICSE 2025 · 7 citations
