Leveraging Automated Unit Tests for Unsupervised Code Translation
Baptiste Rozière, Jie Zhang, François Charton, Mark Harman, Gabriel Synnaeve, Guillaume Lample
Abstract
With little to no parallel data available for programming languages, unsupervised methods are well-suited to source code translation. However, the majority of unsupervised machine translation approaches rely on back-translation, a method developed in the context of natural language translation and one that inherently involves training on noisy inputs. Unfortunately, source code is highly sensitive to small changes; a single token can result in compilation failures or erroneous programs, unlike natural languages where small inaccuracies may not change the meaning of a sentence. To address this issue, we propose to leverage an automated unit-testing system to filter out invalid translations, thereby creating a fully tested parallel corpus. We found that fine-tuning an unsupervised model with this filtered data set significantly reduces the noise in the translations so-generated, comfortably outperforming the state-of-the-art for all language pairs studied. In particular, for Java Python and Python C++ we outperform the best previous methods by more than 16% and 24% respectively, reducing the error rate by more than 35%.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 843ac5c2-3d57-40b9-b855-05b91369242dCited by top-tier papers51
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 2,317 citations
- Teaching Large Language Models to Self-DebugXinyun Chen, Maxwell Lin, Nathanael Schärli, Denny ZhouICLR 2024 · 1,085 citations
- Lost in Translation: A Study of Bugs Introduced by Large Language Models while Translating CodeRangeet Pan, Ali Reza Ibrahimzada, Rahul Krishna, Divya Sankar et al.ICSE 2024 · 96 citations
- Exploring and Unleashing the Power of Large Language Models in Automated Code TranslationZhen Yang, Fang Liu, Zhongxing Yu, Jacky Wai Keung et al.FSE 2024 · 72 citations
- CodeT: Code Generation with Generated TestsBei Chen, Fengji Zhang, Anh Nguyen, Daoguang Zan et al.ICLR 2023 · 64 citations
Builds on7
- GraphCodeBERT: Pre-training Code Representations with Data FlowDaya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng et al.ICLR 2021 · 1,644 citations
- Unsupervised Translation of Programming LanguagesBaptiste Rozière, Marie-Anne Lachaux, Lowik Chanussot, Guillaume LampleNeurIPS 2020 · 606 citations
- Learning and Evaluating Contextual Embedding of Source CodeAditya Kanade, Petros Maniatis, Gogul Balakrishnan, Kensen ShiICML 2020 · 438 citations
- Hoppity: Learning Graph Transformations to Detect and Fix Bugs in ProgramsElizabeth Dinella, Hanjun Dai, Ziyang Li, Mayur Naik et al.ICLR 2020 · 212 citations
- Graph-based, Self-Supervised Program Repair from Diagnostic FeedbackMichihiro Yasunaga, Percy LiangICML 2020 · 198 citations
Related papers
- Syntax and Domain Aware Model for Unsupervised Program TranslationFang Liu, Jia Li, Li ZhangICSE 2023 · 25 citations
- Program Translation via Code DistillationYufan Huang, Mengnan Qi, Yongqiang Yao, Maoquan Wang et al.EMNLP 2023 · 7 citations
- Semi-Supervised Code Translation Overcoming the Scarcity of Parallel Code DataMing Zhu, Mohimenul Karim, Ismini Lourentzou, Daphne YaoASE 2024 · 2 citations
- TransMap: Pinpointing Mistakes in Neural Code TranslationBo Wang, Ruishi Li, Mingkai Li, Prateek SaxenaFSE 2023 · 4 citations
- Code Translation with Compiler RepresentationsMarc Szafraniec, Baptiste Rozière, Hugh Leather, Patrick Labatut et al.ICLR 2023 · 18 citations
