BabelTower: Learning to Auto-parallelized Program Translation
Yuanbo Wen, Qi Guo, Qiang Fu, Xiaqing Li, Jianxing Xu, Yanlin Tang, Yongwei Zhao, Xing Hu, Zidong Du, Ling Li, Chao Wang, Xuehai Zhou, Yunji Chen
Abstract
GPUs have become the dominant computing platforms for many applications, while programming GPUs with the widely-used CUDA parallel programming model is difficult. As sequential C code is relatively easy to obtain either from legacy repositories or by manual implementation, automatically translating C to its parallel CUDA counterpart is promising to relieve the burden of GPU programming. However, because of huge differences between the sequential C and the parallel CUDA programming model, existing approaches fail to conduct the challenging auto-parallelized program translation. In this paper, we propose a learning-based framework, i.e., BabelTower, to address this problem. We first create a large-scale dataset consisting of computeintensive function-level monolingual corpora. We further propose using back-translation with a discriminative reranker to cope with unpaired corpora and parallel semantic conversion. Experimental results show that BabelTower outperforms state-of-the-art by 1.79, 6.09, and 9.39 in terms of BLEU, CodeBLEU, and specifically designed ParaBLEU, respectively. The CUDA code generated by BabelTower attains a speedup of up to 347× over the sequential C code, and the developer productivity is improved by at most 3.8×.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b0782fe9-f22c-48fb-822a-0ce55d65bc18Cited by top-tier papers3
- Lost in Translation: A Study of Bugs Introduced by Large Language Models while Translating CodeRangeet Pan, Ali Reza Ibrahimzada, Rahul Krishna, Divya Sankar et al.ICSE 2024 · 96 citations
- A Joint Learning Model with Variational Interaction for Multilingual Program TranslationYali Du, Hui Sun, Ming LiASE 2024 · 6 citations
- ExeCoder: Empowering Large Language Models with Executability Representation for Code TranslationMinghua He, Yue Chen, Fangkai Yang, Pu Zhao et al.EMNLP 2025 · 1 citation
Builds on4
- Unsupervised Translation of Programming LanguagesBaptiste Rozière, Marie-Anne Lachaux, Lowik Chanussot, Guillaume LampleNeurIPS 2020 · 606 citations
- DOBF: A Deobfuscation Pre-Training Objective for Programming LanguagesMarie-Anne Lachaux, Baptiste Rozière, Marc Szafraniec, Guillaume LampleNeurIPS 2021 · 174 citations
- Leveraging Automated Unit Tests for Unsupervised Code TranslationBaptiste Rozière, Jie Zhang, François Charton, Mark Harman et al.ICLR 2022 · 161 citations
- Discriminative Reranking for Neural Machine TranslationAnn Lee, Michael Auli, Marc'Aurelio RanzatoACL 2021
Related papers
- QiMeng-MuPa: Mutual-Supervised Learning for Sequential-to-Parallel Code TranslationChangxin Ke, Rui Zhang, Shuo Wang, Li Ding et al.NeurIPS 2025 · 3 citations
- CodeRosetta: Pushing the Boundaries of Unsupervised Code Translation for Parallel ProgrammingAli TehraniJamsaz, Arijit Bhattacharjee, Le Chen, Nesreen K. Ahmed et al.NeurIPS 2024 · 36 citations
- Beyond Code Pairs: Dialogue-Based Data Generation for LLM Code TranslationLe Chen, Nuo Xu, Winson Chen, Bin Lei et al.ACL 2026 · 6 citations
- Understanding Performance Problems in CUDA ProgramsYuyang Bi, Junming Cao, You Lu, Bihuan Chen et al.FSE 2026
- QiMeng-Xpiler: Transcompiling Tensor Programs for Deep Learning Systems with a Neural-Symbolic ApproachShouyang Dong, Jun Bi, Di Huang, Jiaming Guo et al.OSDI 2025 · 6 citations
