Scalable, Validated Code Translation of Entire Projects using Large Language Models
Hanliang Zhang, Cristina David, Meng Wang, Brandon Paulsen, Daniel Kroening
Abstract
Large language models (LLMs) show promise in code translation due to their ability to generate idiomatic code. However, a significant limitation when using LLMs for code translation is scalability: existing works have shown a drop in translation success rates for code exceeding around 100 lines. We overcome this limitation by developing a modular approach to translation, where we partition the code into small code fragments which can be translated independently and semantically validated (that is, by checking I/O equivalence). When this approach is applied naively, we discover that LLMs are unreliable when translating features of the source language that do not have a direct mapping to the target language, and that the LLM often gets stuck in repair loops when attempting to fix errors. To address these issues, we introduce two key concepts: (1) feature mapping , which integrates predefined translation rules with LLM-based translation to guide the LLM in navigating subtle language differences and producing semantically accurate code; and (2) type-compatibility , which facilitates localized checks at the function signature level to detect errors early, thereby narrowing the scope of potential repairs. We apply our approach to translating real-world Go codebases to Rust, demonstrating that we can consistently generate reliable Rust translations for projects up to 9,700 lines of code and 780 functions, with an average of 73% of functions successfully validated for I/O equivalence, considerably higher than any existing work.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7d431688-2888-4dd9-b906-5278e95af3a9Cited by top-tier papers11
- MatchFixAgent: Language-Agnostic Autonomous Repository-Level Code Translation Validation and RepairAli Reza Ibrahimzada, Brandon Paulsen, Reyhaneh Jabbarvand, Joey Dodds et al.ICML 2026 · 9 citations
- RustAssure: Differential Symbolic Testing for LLM-Transpiled C-to-Rust CodeYubo Bai, Tapti PalitASE 2025 · 8 citations
- AlphaTrans: A Neuro-Symbolic Compositional Approach for Repository-Level Code Translation and ValidationAli Reza Ibrahimzada, Kaiyao Ke, Mrigank Pawagi, Muhammad Salman Abid et al.FSE 2025 · 8 citations
- SmartC2Rust: Iterative, Feedback-Driven C-to-Rust Translation via Large Language Models for Safety and EquivalenceMomoko Shiraishi, Yinzhi Cao, Takahiro ShinagawaICSE 2026 · 4 citations
- RustRepoTrans: Repository-level Context Code Translation Benchmark Targeting RustGuangsheng Ou, Mingwei Liu, Yuxuan Chen, Yanlin Wang et al.ASE 2025 · 1 citation
Builds on15
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 2,317 citations
- Unsupervised Translation of Programming LanguagesBaptiste Rozière, Marie-Anne Lachaux, Lowik Chanussot, Guillaume LampleNeurIPS 2020 · 606 citations
- Automated Program Repair in the Era of Large Pre-trained Language ModelsChunqiu Steven Xia, Yuxiang Wei, Lingming ZhangICSE 2023 · 321 citations
- Leveraging Automated Unit Tests for Unsupervised Code TranslationBaptiste Rozière, Jie Zhang, François Charton, Mark Harman et al.ICLR 2022 · 161 citations
- NEZHA: Efficient Domain-Independent Differential TestingTheofilos Petsios, Adrian Tang, Salvatore J. Stolfo, Angelos D. Keromytis et al.S&P 2017 · 132 citations
Related papers
- SACTOR: LLM-Driven Correct and Idiomatic C to Rust Translation with Static Analysis and FFI-Based VerificationTianyang Zhou, Ziyi Zhang, Haowen Lin, Somesh Jha et al.ACL 2026 · 8 citations
- VERT: Polyglot Verified Equivalent Rust Transpilation with Large Language ModelsAidan Z. H. Yang, Yoshiki Takashima, Brandon Paulsen, Josiah Dodds et al.ASE 2025 · 1 citation
- Lost in Translation: A Study of Bugs Introduced by Large Language Models while Translating CodeRangeet Pan, Ali Reza Ibrahimzada, Rahul Krishna, Divya Sankar et al.ICSE 2024 · 96 citations
- Code Translation with Compiler RepresentationsMarc Szafraniec, Baptiste Rozière, Hugh Leather, Patrick Labatut et al.ICLR 2023 · 18 citations
- RustAssistant: Using LLMs to Fix Compilation Errors in Rust CodePantazis Deligiannis, Akash Lal, Nikita Mehrotra, Rishi Poddar et al.ICSE 2025 · 6 citations
