An interpretable error correction method for enhancing code-to-code translation
Min Xue, Artur Andrzejak, Marla Leuther
Abstract
Transformer-based machine translation models currently dominate the field of model-based program translation. However, these models fail to provide interpretative support for the generated program translations. Moreover, researchers frequently invest substantial time and computational resources in retraining models, yet the improvement in translation accuracy is quite limited. To address these issues, we introduce a novel approach, kNN-ECD, which combines k-nearestneighbor search with a key-value error correction datastore to overwrite the wrong translations of TransCoder-ST (Roziere et al., 2022). This provides a decisionmaking basis for interpreting the corrected translations. Building upon this, we further propose kNN-ECS m , a methodology that employs a distributed structure with m sub-datastores connected in series, utilizing m diverse experts for multiround error correction. Additionally, we put forward a unified name rule, encouraging the datastore to focus more on code logic and structure rather than diverse rare identifiers. Our experimental results show that our approach improves the translation accuracy from 68.9% to 89.9% of TransCoder-ST (for translation from Java to Python). This error correction method augments program translation, overcoming the inherent limitations of Transformer-based code translation models, such as resource-intensive retraining requirements and uninterpretable outcomes. * Corresponding author. * ⟨•, •⟩ → '* ' : ⟨•, •⟩ represents the decision-making basis for each generated token '* ' in the corrected Python function. * NEW LINE, INDENT, DEDENT represent newline, indentation, dedent in code formatting, respectively.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext eaa4bf39-0bd7-4e7e-ae10-ec6569a5214cCited by top-tier papers2
- SmartC2Rust: Iterative, Feedback-Driven C-to-Rust Translation via Large Language Models for Safety and EquivalenceMomoko Shiraishi, Yinzhi Cao, Takahiro ShinagawaICSE 2026 · 4 citations
- Function-to-Style Guidance of LLMs for Code TranslationLonghui Zhang, Bin Wang, Jiahao Wang, Xiaofeng Zhao et al.ICML 2025
Builds on17
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 3,415 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- Investigating Gender Bias in Language Models Using Causal Mediation AnalysisJesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian et al.NeurIPS 2020 · 851 citations
- Unsupervised Translation of Programming LanguagesBaptiste Rozière, Marie-Anne Lachaux, Lowik Chanussot, Guillaume LampleNeurIPS 2020 · 606 citations
- Nearest Neighbor Machine TranslationUrvashi Khandelwal, Angela Fan, Dan Jurafsky, Luke Zettlemoyer et al.ICLR 2021 · 323 citations
Related papers
- Precise and Interpretable Editing of Code Knowledge in Large Language ModelsMin Xue, Nikolai Bolik, Lennart Stöpler, Erik Imgrund et al.ICLR 2026
- Learning to Recommend Method Names with Global ContextFang Liu, Ge Li, Zhiyi Fu, Shuai Lu et al.ICSE 2022 · 32 citations
- TransMap: Pinpointing Mistakes in Neural Code TranslationBo Wang, Ruishi Li, Mingkai Li, Prateek SaxenaFSE 2023 · 4 citations
- Bridging the Domain Gaps in Context Representations for k-Nearest Neighbor Neural Machine TranslationZhiwei Cao, Baosong Yang, Huan Lin, Suhang Wu et al.ACL 2023 · 1 citation
- Towards Robust k-Nearest-Neighbor Machine TranslationHui Jiang, Ziyao Lu, Fandong Meng, Chulun Zhou et al.EMNLP 2022 · 16 citations
