Scalable, Validated Code Translation of Entire Projects using Large Language Models
Hanliang Zhang, Cristina David, Meng Wang, Brandon Paulsen, Daniel Kroening
摘要
Large language models (LLMs) show promise in code translation due to their ability to generate idiomatic code. However, a significant limitation when using LLMs for code translation is scalability: existing works have shown a drop in translation success rates for code exceeding around 100 lines. We overcome this limitation by developing a modular approach to translation, where we partition the code into small code fragments which can be translated independently and semantically validated (that is, by checking I/O equivalence). When this approach is applied naively, we discover that LLMs are unreliable when translating features of the source language that do not have a direct mapping to the target language, and that the LLM often gets stuck in repair loops when attempting to fix errors. To address these issues, we introduce two key concepts: (1) feature mapping , which integrates predefined translation rules with LLM-based translation to guide the LLM in navigating subtle language differences and producing semantically accurate code; and (2) type-compatibility , which facilitates localized checks at the function signature level to detect errors early, thereby narrowing the scope of potential repairs. We apply our approach to translating real-world Go codebases to Rust, demonstrating that we can consistently generate reliable Rust translations for projects up to 9,700 lines of code and 780 functions, with an average of 73% of functions successfully validated for I/O equivalence, considerably higher than any existing work.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- MatchFixAgent: Language-Agnostic Autonomous Repository-Level Code Translation Validation and RepairAli Reza Ibrahimzada, Brandon Paulsen, Reyhaneh Jabbarvand, Joey Dodds 等ICML 2026 · 被引用 9 次
- RustAssure: Differential Symbolic Testing for LLM-Transpiled C-to-Rust CodeYubo Bai, Tapti PalitASE 2025 · 被引用 8 次
- AlphaTrans: A Neuro-Symbolic Compositional Approach for Repository-Level Code Translation and ValidationAli Reza Ibrahimzada, Kaiyao Ke, Mrigank Pawagi, Muhammad Salman Abid 等FSE 2025 · 被引用 8 次
- SmartC2Rust: Iterative, Feedback-Driven C-to-Rust Translation via Large Language Models for Safety and EquivalenceMomoko Shiraishi, Yinzhi Cao, Takahiro ShinagawaICSE 2026 · 被引用 4 次
- RustRepoTrans: Repository-level Context Code Translation Benchmark Targeting RustGuangsheng Ou, Mingwei Liu, Yuxuan Chen, Yanlin Wang 等ASE 2025 · 被引用 1 次
它引用的顶会 Paper15
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 被引用 2,317 次
- Unsupervised Translation of Programming LanguagesBaptiste Rozière, Marie-Anne Lachaux, Lowik Chanussot, Guillaume LampleNeurIPS 2020 · 被引用 606 次
- Automated Program Repair in the Era of Large Pre-trained Language ModelsChunqiu Steven Xia, Yuxiang Wei, Lingming ZhangICSE 2023 · 被引用 321 次
- Leveraging Automated Unit Tests for Unsupervised Code TranslationBaptiste Rozière, Jie Zhang, François Charton, Mark Harman 等ICLR 2022 · 被引用 161 次
- NEZHA: Efficient Domain-Independent Differential TestingTheofilos Petsios, Adrian Tang, Salvatore J. Stolfo, Angelos D. Keromytis 等S&P 2017 · 被引用 132 次
相关 Paper
- SACTOR: LLM-Driven Correct and Idiomatic C to Rust Translation with Static Analysis and FFI-Based VerificationTianyang Zhou, Ziyi Zhang, Haowen Lin, Somesh Jha 等ACL 2026 · 被引用 8 次
- VERT: Polyglot Verified Equivalent Rust Transpilation with Large Language ModelsAidan Z. H. Yang, Yoshiki Takashima, Brandon Paulsen, Josiah Dodds 等ASE 2025 · 被引用 1 次
- Lost in Translation: A Study of Bugs Introduced by Large Language Models while Translating CodeRangeet Pan, Ali Reza Ibrahimzada, Rahul Krishna, Divya Sankar 等ICSE 2024 · 被引用 96 次
- Code Translation with Compiler RepresentationsMarc Szafraniec, Baptiste Rozière, Hugh Leather, Patrick Labatut 等ICLR 2023 · 被引用 18 次
- RustAssistant: Using LLMs to Fix Compilation Errors in Rust CodePantazis Deligiannis, Akash Lal, Nikita Mehrotra, Rishi Poddar 等ICSE 2025 · 被引用 6 次
