Code Translation with Compiler Representations
Marc Szafraniec, Baptiste Rozière, Hugh Leather, Patrick Labatut, François Charton, Gabriel Synnaeve
摘要
In this paper, we leverage low-level compiler intermediate representations (IR) to improve code translation. Traditional transpilers rely on syntactic information and handcrafted rules, which limits their applicability and produces unnatural-looking code. Applying neural machine translation (NMT) approaches to code has successfully broadened the set of programs on which one can get a natural-looking translation. However, they treat the code as sequences of text tokens, and still do not differentiate well enough between similar pieces of code which have different semantics in different languages. The consequence is low quality translation, reducing the practicality of NMT, and stressing the need for an approach significantly increasing its accuracy. Here we propose to augment code translation with IRs, specifically LLVM IR, with results on the C++, Java, Rust, and Go languages. Our method improves upon the state of the art for unsupervised code translation, increasing the number of correct translations by 11% on average, and up to 79% for the Java -> Rust pair with greedy decoding. We extend previous test sets for code translation, by adding hundreds of Go and Rust functions. Additionally, we train models with high performance on the problem of IR decompilation, generating programming source code from IR, and study using IRs as intermediary pivot for translation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper28
- Lost in Translation: A Study of Bugs Introduced by Large Language Models while Translating CodeRangeet Pan, Ali Reza Ibrahimzada, Rahul Krishna, Divya Sankar 等ICSE 2024 · 被引用 96 次
- Exploring and Unleashing the Power of Large Language Models in Automated Code TranslationZhen Yang, Fang Liu, Zhongxing Yu, Jacky Wai Keung 等FSE 2024 · 被引用 72 次
- Multi-lingual Evaluation of Code Generation ModelsBen Athiwaratkun, Sanjay Krishna Gouda, Zijian Wang, Xiaopeng Li 等ICLR 2023 · 被引用 28 次
- Scalable, Validated Code Translation of Entire Projects using Large Language ModelsHanliang Zhang, Cristina David, Meng Wang, Brandon Paulsen 等PLDI 2025 · 被引用 15 次
- GALLa: Graph Aligned Large Language Models for Improved Source Code UnderstandingZiyin Zhang, Hang Yu, Sage Lee, Peng Di 等ACL 2025 · 被引用 11 次
它引用的顶会 Paper4
- GraphCodeBERT: Pre-training Code Representations with Data FlowDaya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng 等ICLR 2021 · 被引用 1,644 次
- Unsupervised Translation of Programming LanguagesBaptiste Rozière, Marie-Anne Lachaux, Lowik Chanussot, Guillaume LampleNeurIPS 2020 · 被引用 606 次
- Learning and Evaluating Contextual Embedding of Source CodeAditya Kanade, Petros Maniatis, Gogul Balakrishnan, Kensen ShiICML 2020 · 被引用 438 次
- Leveraging Automated Unit Tests for Unsupervised Code TranslationBaptiste Rozière, Jie Zhang, François Charton, Mark Harman 等ICLR 2022 · 被引用 161 次
相关 Paper
- Program Translation via Code DistillationYufan Huang, Mengnan Qi, Yongqiang Yao, Maoquan Wang 等EMNLP 2023 · 被引用 7 次
- VERT: Polyglot Verified Equivalent Rust Transpilation with Large Language ModelsAidan Z. H. Yang, Yoshiki Takashima, Brandon Paulsen, Josiah Dodds 等ASE 2025 · 被引用 1 次
- IRCoder: Intermediate Representations Make Language Models Robust Multilingual Code GeneratorsIndraneil Paul, Goran Glavas, Iryna GurevychACL 2024
- RustRepoTrans: Repository-level Context Code Translation Benchmark Targeting RustGuangsheng Ou, Mingwei Liu, Yuxuan Chen, Yanlin Wang 等ASE 2025 · 被引用 1 次
- Semi-Supervised Code Translation Overcoming the Scarcity of Parallel Code DataMing Zhu, Mohimenul Karim, Ismini Lourentzou, Daphne YaoASE 2024 · 被引用 2 次
