On-Policy Optimization with Group Equivalent Preference for Multi-Programming Language Understanding
Haoyuan Wu, Rui Ming, Jilong Gao, Hangyu Zhao, Xueyi Chen, Yikai Yang, Haisheng Zheng, Zhuolun He, Bei Yu
Abstract
Large language models (LLMs) achieve remarkable performance in code generation tasks. However, a significant performance disparity persists between popular programming languages (e.g., Python, C++) and others. To address this capability gap, we leverage the code translation task to train LLMs, thereby facilitating the transfer of coding proficiency across diverse programming languages. Moreover, we introduce OORL for training, a novel reinforcement learning (RL) framework that integrates on-policy and off-policy strategies. Within OORL, on-policy RL is applied during code translation, guided by a rule-based reward signal derived from unit tests. Complementing this coarse-grained rule-based reward, we propose Group Equivalent Preference Optimization (GEPO), a novel preference optimization method. Specifically, GEPO trains the LLM using intermediate representations (IRs) groups. LLMs can be guided to discern IRs equivalent to the source code from inequivalent ones, while also utilizing signals about the mutual equivalence between IRs within the group. This process allows LLMs to capture nuanced aspects of code functionality. By employing OORL for training with code translation tasks, LLMs improve their recognition of code functionality and their understanding of the relationships between code implemented in different languages. Extensive experiments demonstrate that our OORL for LLMs training with code translation tasks achieves significant performance improvements on code benchmarks across multiple programming languages.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 37d01fe9-e985-4702-9983-b6735c2c6583Cited by top-tier papers2
- StreamingTOM: Streaming Token Compression for Efficient Video UnderstandingXueyi Chen, Keda Tao, Kele Shao, Huan WangCVPR 2026 · 46 citations
- Bootstrapping Code Translation with Weighted Multilanguage ExplorationYuhan Wu, Huan Zhang, Wei Cheng, Chen Shen et al.ACL 2026
Builds on6
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 852 citations
- ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language ModelsZiniu Li, Tian Xu, Yushun Zhang, Zhihang Lin et al.ICML 2024 · 165 citations
- Back to Basics: Revisiting REINFORCE-Style Optimization for Learning from Human Feedback in LLMsArash Ahmadian, Chris Cremer, Matthias Gallé, Marzieh Fadaee et al.ACL 2024 · 20 citations
- Code Translation with Compiler RepresentationsMarc Szafraniec, Baptiste Rozière, Hugh Leather, Patrick Labatut et al.ICLR 2023 · 18 citations
- Program Translation via Code DistillationYufan Huang, Mengnan Qi, Yongqiang Yao, Maoquan Wang et al.EMNLP 2023 · 7 citations
Related papers
- GXPO: Group Cross-Lingual Relative Policy Optimization for Code GenerationLinzheng Chai, Jian Yang, Jiajun Wu, Ensheng Shi et al.ICML 2026
- Empowering Multi-Turn Tool-Integrated Agentic Reasoning with Group Turn Policy OptimizationYifeng Ding, Hung Le, Songyang Han, Kangrui Ruan et al.ACL 2026 · 5 citations
- ReCode: Updating Code API Knowledge with Reinforcement LearningHaoze Wu, Yunzhi Yao, Wenhao Yu, Ningyu ZhangAAAI 2026 · 7 citations
- Lost in Translation: A Study of Bugs Introduced by Large Language Models while Translating CodeRangeet Pan, Ali Reza Ibrahimzada, Rahul Krishna, Divya Sankar et al.ICSE 2024 · 96 citations
- Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency OptimizationMingzhe Du, Anh Tuan Luu, Yue Liu, Yuhao Qing et al.NeurIPS 2025 · 18 citations
