Learning LLM-as-a-Judge for Preference Alignment
Ziyi Ye, Xiangsheng Li, Qiuchi Li, Qingyao Ai, Yujia Zhou, Wei Shen, Dong Yan, Yiqun Liu
Abstract
Learning from preference feedback is a common practice for aligning large language models (LLMs) with human value. Conventionally, preference data is learned and encoded into a scalar reward model that connects a value head with an LLM to produce a scalar score as preference. However, scalar models lack interpretability and are known to be susceptible to biases in datasets. This paper investigates leveraging LLM itself to learn from such preference data and serve as a judge to address both limitations in one shot. Specifically, we prompt the pre-trained LLM to generate initial judgment pairs with contrastive preference in natural language form. The self-generated contrastive judgment pairs are used to train the LLM-as-a-Judge with Direct Preference Optimization (DPO) and incentivize its reasoning capability as a judge. This proposal of learning the LLMas-a-Judge using self-generated Contrastive judgments (Con-J) ensures natural interpretability through the generated rationales supporting the judgments, and demonstrates higher robustness against bias compared to scalar models. Experimental results show that Con-J outperforms the scalar reward model trained on the same collection of preference data, and outperforms a series of open-source and closed-source generative LLMs. We open-source the training process and model weights of Con-J at https://github.com/YeZiyi1998/Con-J.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a799f40e-08fb-4ba4-8c02-cf26a84727afCited by top-tier papers10
- Reward Reasoning ModelsJiaxin Guo, Zewen Chi, Li Dong, Qingxiu Dong et al.NeurIPS 2025 · 14 citations
- Sparta Alignment: Collectively Aligning Multiple Language Models through CombatYuru Jiang, Wenxuan Ding, Shangbin Feng, Greg Durrett et al.NeurIPS 2025 · 8 citations
- P-Check: Advancing Personalized Reward Model via Learning to Generate Dynamic ChecklistKwangwook Seo, Dongha LeeACL 2026 · 2 citations
- SimVBG: Simulating Individual Values by Backstory GenerationBangde Du, Ziyi Ye, Zhijing Wu, Monika Jankowska et al.EMNLP 2025 · 2 citations
- Distribution-Calibrated Inference Time Compute for Thinking LLM-as-a-JudgeHamid Dadkhahi, Firas Trabelsi, Parker Riley, Juraj Juraska et al.ICML 2026 · 2 citations
Builds on18
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- STaR: Bootstrapping Reasoning With ReasoningEric Zelikman, Yuhuai Wu, Jesse Mu, Noah D. GoodmanNeurIPS 2022 · 1,126 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 852 citations
Related papers
- Cal-DPO: Calibrated Direct Preference Optimization for Language Model AlignmentTeng Xiao, Yige Yuan, Huaisheng Zhu, Mingxiao Li et al.NeurIPS 2024 · 76 citations
- Direct Judgement Preference OptimizationPeifeng Wang, Austin Xu, Yilun Zhou, Caiming Xiong et al.EMNLP 2025 · 1 citation
- Improve LLM-as-a-Judge Ability as a General AbilityJiachen Yu, Shaoning Sun, Xiaohui Hu, Jiaxu Yan et al.EMNLP 2025 · 1 citation
- Direct Large Language Model Alignment Through Self-Rewarding Contrastive Prompt DistillationAiwei Liu, Haoping Bai, Zhiyun Lu, Xiang Kong et al.ACL 2024 · 4 citations
- What Matters in Data for DPO?Yu Pan, Zhongze Cai, Huaiyang Zhong, Guanting Chen et al.NeurIPS 2025 · 13 citations
