Bootstrapping Code Translation with Weighted Multilanguage Exploration
Yuhan Wu, Huan Zhang, Wei Cheng, Chen Shen, Jingyue Yang, Wei Hu
Abstract
Code translation across multiple programming languages is essential yet challenging due to two vital obstacles: scarcity of parallel data paired with executable test oracles, and optimization imbalance when handling diverse language pairs. We propose BootTrans, a bootstrapping method that resolves both obstacles. Its key idea is to leverage the functional invariance and cross-lingual portability of test suites, adapting abundant pivot-language unit tests to serve as universal verification oracles for multilingual reinforcement learning (RL) training. Our method introduces a dual-pool architecture with seed and exploration pools to progressively expand training data via execution-guided experience collection. Furthermore, we design a language-aware weighting mechanism that dynamically prioritizes harder translation directions based on relative performance across sibling languages, mitigating optimization imbalance. Extensive experiments on the HumanEval-X and TransCoder-Test benchmarks demonstrate substantial improvements over baseline LLMs across all translation directions, with ablation studies validating the effectiveness of both bootstrapping and weighting components.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2a269b95-cdc7-43e5-a563-bc11cc7e1f0fBuilds on18
- GraphCodeBERT: Pre-training Code Representations with Data FlowDaya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng et al.ICLR 2021 · 1,644 citations
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 1,224 citations
- Unsupervised Translation of Programming LanguagesBaptiste Rozière, Marie-Anne Lachaux, Lowik Chanussot, Guillaume LampleNeurIPS 2020 · 606 citations
- CodeRL: Mastering Code Generation through Pretrained Models and Deep Reinforcement LearningHung Le, Yue Wang, Akhilesh Deepak Gotmare, Silvio Savarese et al.NeurIPS 2022 · 571 citations
- Leveraging Automated Unit Tests for Unsupervised Code TranslationBaptiste Rozière, Jie Zhang, François Charton, Mark Harman et al.ICLR 2022 · 161 citations
Related papers
- On-Policy Optimization with Group Equivalent Preference for Multi-Programming Language UnderstandingHaoyuan Wu, Rui Ming, Jilong Gao, Hangyu Zhao et al.NeurIPS 2025 · 2 citations
- Bilingual alignment transfers to multilingual alignment for unsupervised parallel text miningChih-chan Tien, Shane Steinert-ThrelkeldACL 2022 · 10 citations
- GXPO: Group Cross-Lingual Relative Policy Optimization for Code GenerationLinzheng Chai, Jian Yang, Jiajun Wu, Ensheng Shi et al.ICML 2026
- INTERTRANS: Leveraging Transitive Intermediate Translations to Enhance LLM-Based Code TranslationMarcos Macedo, Yuan Tian, Pengyu Nie, Filipe Roseiro Côgo et al.ICSE 2025 · 7 citations
- Balancing Training for Multilingual Neural Machine TranslationXinyi Wang, Yulia Tsvetkov, Graham NeubigACL 2020 · 74 citations
