Direct Judgement Preference Optimization
Peifeng Wang, Austin Xu, Yilun Zhou, Caiming Xiong, Shafiq Joty
Abstract
To meet the increasing need for timely and accurate evaluation of large language model (LLM) responses, training LLM-as-judges to evaluate and critique other model responses has emerged as a popular paradigm. However, existing judge models are largely trained with supervised finetuning (SFT) on small data scales to perform limited types of evaluation tasks, fundamentally limiting generalization. To meet the need for strong, generalized judge models, we explore training foundational judge models at large data scales (680K) with direct preference optimization (DPO). Using four training tasks, we form three types of DPO preference pairs targeting different aspects of evaluation: Generating meaningful critiques, making accurate judgements, and understanding what comprises good and bad responses. To demonstrate the effectiveness of our method, we train judge models of three sizes: 8B parameters, 12B, and 70B, and evaluate on a comprehensive suite of 13 benchmarks (7 pairwise, 4 single rating, and 2 classification). Our models achieve the best aggregate performance, with even our 8B model outperforming GPT-4o in pairwise benchmarks. Further analysis shows that our judge models produce factual and actionable critiques and serve as strong foundational judges for continued finetuning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d778f249-24e3-4c6f-b975-715fa963f59aCited by top-tier papers7
- Skywork-Reward-V2: Scaling Preference Data Curation via Human-AI SynergyChris Yuhao Liu, Liang Zeng, Yuzhen Xiao, Jujie He et al.ICLR 2026 · 211 citations
- J1: Incentivizing Thinking in LLM-as-a-Judge via Reinforcement LearningChenxi Whitehouse, Tianlu Wang, Ping Yu, Xian Li et al.ICLR 2026 · 74 citations
- Variation in Verification: Understanding Verification Dynamics in Large Language ModelsYefan Zhou, Austin Xu, Yilun Zhou, Janvijay Singh et al.ICLR 2026 · 17 citations
- IF-CRITIC: Towards a Fine-Grained LLM Critic for Instruction-Following EvaluationBosi Wen, Yilin Niu, Cunxiang Wang, Pei Ke et al.ACL 2026 · 2 citations
- Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-JudgeSwarnadeep Saha, Xian Li, Marjan Ghazvininejad, Jason E. Weston et al.ICML 2025
Builds on28
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- Multitask Prompted Training Enables Zero-Shot Task GeneralizationVictor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach et al.ICLR 2022 · 1,976 citations
- SimPO: Simple Preference Optimization with a Reference-Free RewardYu Meng, Mengzhou Xia, Danqi ChenNeurIPS 2024 · 1,203 citations
Related papers
- Improve LLM-as-a-Judge Ability as a General AbilityJiachen Yu, Shaoning Sun, Xiaohui Hu, Jiaxu Yan et al.EMNLP 2025 · 1 citation
- Learning LLM-as-a-Judge for Preference AlignmentZiyi Ye, Xiangsheng Li, Qiuchi Li, Qingyao Ai et al.ICLR 2025
- CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding TasksHongchao Jiang, Yiming Chen, Yushi Cao, Hung-Yi Lee et al.ACL 2026 · 33 citations
- JudgeBench: A Benchmark for Evaluating LLM-Based JudgesSijun Tan, Siyuan Zhuang, Kyle Montgomery, William Yuan Tang et al.ICLR 2025
- Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling EvaluatorsYilun Zhou, Austin Xu, Peifeng Wang, Caiming Xiong et al.ICML 2025
