Lune

ASE2025顶会

Polyglot: An Extensible Framework to Benchmark Code Translation with LLMs

Marco Vieira, Priyam Ashish Shah, Bhavain Shah, Rrezarta Krasniqi

2025年份

摘要

Large Language Models (LLMs) show great potential for automating code-related tasks. However, sound assessments are necessary to understand their true capabilities, particularly in code translation, where reliability is crucial. We introduce Polyglot, an automated, multi-language framework for evaluating the translation quality of LLMs between different programming languages. Leveraging the IBM CodeNet Project, an extensive collection of coding problems in multiple languages, we assess translation quality using syntactic correctness, execution reliability, semantic preservation, and static code metrics. Our evaluation focuses on translating C to Java, Python, and Rust, languages that follow distinct paradigms and represent alternatives to modernize C-based systems. We evaluate open-source LLMs using three prompting strategies to understand the impact on translation performance. Our findings highlight that while LLMs show promising results for simple code translation, their limitations regarding complex logic and distinct language paradigms require further analysis.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

lune papers get 89f63c4b-5fe1-4b4e-8293-9002a13e5dee

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖