CodeDPO: Aligning Code Models with Self Generated and Verified Source Code
Kechi Zhang, Ge Li, Yihong Dong, Jingjing Xu, Jun Zhang, Jing Su, Yongfei Liu, Zhi Jin
摘要
Code generation models have shown significant potential for programming tasks. However, existing training methods like supervised fine-tuning face key limitations: they do not effectively teach models to prioritize correct over incorrect solutions in ambiguous situations, nor do they effectively optimize the runtime efficiency of the generated code. To address these challenges, we propose CodeDPO, a framework that integrates preference learning into code generation to improve two key code preference factors: code correctness and efficiency. CodeDPO employs a novel dataset construction method, utilizing a self-generation-andvalidation mechanism that simultaneously generates and evaluates code and test cases. The underlying assumption is that test cases executable by multiple code snippets provide more reliable validation, and code that passes more tests is more likely to be correct. Through this self-validation process, our PageRank-inspired algorithm iteratively updates the ranking score of each code snippet, ultimately creating a code preference optimization dataset based on correctness and efficiency. CodeDPO is flexible and scalable, generating diverse preference optimization data without depending on powerful models such as GPT-4. Through comprehensive evaluations of five widely used benchmarks, CodeDPO demonstrates significant improvements in correctness and efficiency compared to existing methods. Our experiments prove that CodeDPO enhances the capabilities of LLMs in code generation and provides a robust foundation for conducting code preference optimization in more complex and challenging real-world scenarios. 1 Figure 1 : Log probabilities for code with varying correctness and efficiency during Phi-2-2.7B model training on our constructed dataset. The traditional SFT strategy struggles to teach models to prefer correct solutions over incorrect or slow ones. In contrast, our CodeDPO approach effectively optimizes for both correctness and efficiency.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper24
- Visual-RFT: Visual Reinforcement Fine-TuningZiyu Liu, Zeyi Sun, Yuhang Zang, Xiaoyi Dong 等ICCV 2025 · 被引用 563 次
- JarvisArt: Liberating Human Artistic Creativity via an Intelligent Photo Retouching AgentYunlong Lin, Zixu Lin, Kunjie Lin, Jinbin Bai 等NeurIPS 2025 · 被引用 42 次
- GRE Suite: Geo-localization Inference via Fine-Tuned Vision-Language Models and Enhanced Reasoning ChainsChun Wang, Xiaojun Ye, Xiaoran Pan, Zihao Pan 等NeurIPS 2025 · 被引用 18 次
- Fine-R1: Make Multi-modal LLMs Excel in Fine-Grained Visual Recognition by Chain-of-Thought ReasoningHulingxiao He, Zijun Geng, Yuxin PengICLR 2026 · 被引用 12 次
- PEACE: Towards Efficient Project-Level Efficiency Optimization via Hybrid Code EditingXiaoxue Ren, Jun Wan, Yun Peng, Zhongxin Liu 等ASE 2025 · 被引用 5 次
它引用的顶会 Paper11
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Teaching Large Language Models to Self-DebugXinyun Chen, Maxwell Lin, Nathanael Schärli, Denny ZhouICLR 2024 · 被引用 1,085 次
- WizardCoder: Empowering Code Large Language Models with Evol-InstructZiyang Luo, Can Xu, Pu Zhao, Qingfeng Sun 等ICLR 2024 · 被引用 945 次
- Self-Instruct: Aligning Language Models with Self-Generated InstructionsYizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu 等ACL 2023 · 被引用 540 次
- DS-1000: A Natural and Reliable Benchmark for Data Science Code GenerationYuhang Lai, Chengxi Li, Yiming Wang, Tianyi Zhang 等ICML 2023 · 被引用 504 次
相关 Paper
- Teaching Your Models to Understand Code via Focal Preference AlignmentJie Wu, Haoling Li, Xin Zhang, Xiao Liu 等EMNLP 2025
- Alignment with Fill-In-the-Middle for Enhancing Code GenerationHouxing Ren, Zimu Lu, Weikang Shi, Haotian Hou 等EMNLP 2025
- Preference Optimization for Reasoning with Pseudo FeedbackFangkai Jiao, Geyang Guo, Xingxing Zhang, Nancy F. Chen 等ICLR 2025
- Teaching an Old LLM Secure Coding: Localized Preference Optimization on Distilled PreferencesMohammad Saqib Hasan, Saikat Chakraborty, Santu Karmaker, Niranjan BalasubramanianACL 2025
- Process-Supervised Reinforcement Learning for Code GenerationYufan Ye, Ting Zhang, Wenbin Jiang, Hua HuangEMNLP 2025 · 被引用 1 次
