Reward-Free Code Alignment from Pretrained or Fine-Tuned LLM: Unpacking the Trade-offs for Code Generation
Sanjeepan Sivapiran, Gias Uddin
摘要
Large Language Model (LLM) alignment trains an LLM using preference data to produce outputs that better meet established quality standards (e.g., to mitigate toxic or biased responses in text generation, or to enforce coding best practices in code generation). While LLM alignment techniques are studied for non-coding tasks, we know little about their usefulness for coding tasks. Intuitively, for code generation tasks using LLMs, alignment techniques could help with coding solutions that better adhere to established software engineering standards, such as more secure code, supporting better coding practices, etc. However, it is unclear whether LLM code alignment could support both functional requirements (producing executable, correct code) and non-functional requirements (code readability, style, maintainability). It is also unknown whether alignment for a code LLM should begin with base pretrained version or the finetuned (i.e., instruction-tuned) version of the LLM. In this paper, we offer insights on the above two research questions by conducting an empirical study. We studied five state-of-the-art (SOTA) LLMs using two widely used LLM alignment techniques: Direct Preference Optimization (DPO) and BoNBoN. For each training record, we created a preference pair as accepted and rejected instances by using the SelfCodeAlign pipeline. DPO and BoNBoN are reward-free models, i.e., they eliminate the need for multiple reward scores for output preferences. We tuned each LLM using the two alignment techniques in two settings: pretrained and finetuned versions of an LLM. We evaluated functional requirements using four SOTA benchmarks (HumanEval+, MBPP+, EvalPerf, EvoEval) and non-functional requirements using the CODAL benchmark, which evaluates code quality across five dimensions derived from software engineering practices. We find that pretrained-to-aligned pathways achieve larger improvements in the aligned variant over its pretrained variant (CodeLlama-7b: +75% non-functional, Llama3-8b: +42% functional). But the pretrained variant is generally less accurate than its finetuned variant. However, finetuned-to-aligned offers smaller performance improvements or, in some cases, degradation in the aligned variant than its finetuned variant. This means that while the base pretrained version is less accurate than its base finetuned variant, alignment reduces the performance gap between pretrained and finetuned variants. Non-functional requirements improve more consistently than functional requirements via alignment. Based on these findings, we provide nine recommendations to guide alignment for code LLMs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper21
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 被引用 2,317 次
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao 等ICLR 2024 · 被引用 2,082 次
- Model Alignment as Prospect Theoretic OptimizationKawin Ethayarajh, Winnie Xu, Niklas Muennighoff, Dan Jurafsky 等ICML 2024 · 被引用 973 次
相关 Paper
- SelfCodeAlign: Self-Alignment for Code GenerationYuxiang Wei, Federico Cassano, Jiawei Liu, Yifeng Ding 等NeurIPS 2024 · 被引用 79 次
- BoNBoN Alignment for Large Language Models and the Sweetness of Best-of-n SamplingLin Gui, Cristina Garbacea, Victor VeitchNeurIPS 2024 · 被引用 138 次
- Direct Large Language Model Alignment Through Self-Rewarding Contrastive Prompt DistillationAiwei Liu, Haoping Bai, Zhiyun Lu, Xiang Kong 等ACL 2024 · 被引用 4 次
- Test-Time Preference Optimization: On-the-Fly Alignment via Iterative Textual FeedbackYafu Li, Xuyang Hu, Xiaoye Qu, Linjie Li 等ICML 2025
- Dynamic Rewarding with Prompt Optimization Enables Tuning-free Self-Alignment of Language ModelsSomanshu Singla, Zhen Wang, Tianyang Liu, Abdullah Ashfaq 等EMNLP 2024 · 被引用 1 次
