Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge
Swarnadeep Saha, Xian Li, Marjan Ghazvininejad, Jason E. Weston, Tianlu Wang
Abstract
LLM-as-a-Judge models generate chain-ofthought (CoT) sequences intended to capture the step-by-step reasoning process that underlies the final evaluation of a response. However, due to the lack of human-annotated CoTs for evaluation, the required components and structure of effective reasoning traces remain understudied. Consequently, previous approaches often (1) constrain reasoning traces to hand-designed components, such as a list of criteria, reference answers, or verification questions and (2) structure them such that planning is intertwined with the reasoning for evaluation. In this work, we propose EvalPlanner, a preference optimization algorithm for Thinking-LLM-as-a-Judge that first generates an unconstrained evaluation plan, followed by its execution, and then the final judgment. In a self-training loop, EvalPlanner iteratively optimizes over synthetically constructed evaluation plans and executions, leading to better final verdicts. Our method achieves a new state-of-the-art performance for generative reward models on RewardBench and PPE, despite being trained on fewer amount of, and synthetically generated, preference pairs. Additional experiments on other benchmarks like RM-Bench, JudgeBench, and FollowBenchEval further highlight the utility of both planning and reasoning for building robust LLM-as-a-Judge reasoning models. Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge Example Input and Output of EvalPlanner User Instruction: [Text]There is a French pub called The Phoenix that has a low customer rating but is child friendly and is more than 30 dollars and is located in the riverside area near Crowne Plaza Hotel. [Question]According to the above [Text], extract the attributes (name, eatType) of the restaurant. Do not output your analysis, only output the result in the format of "attribute1[value1], attribute2[value2], ..." Response A: Name[The Phoenix], eatType[pub], location[Riverside area near Crowne Plaza Hotel], priceRange[More than 30 dollars] Response B: Name[The Phoenix], eatType[pub] Output: [Start of Evaluation Plan] To evaluate the response quality, we will follow these steps: 1. Check if the response provides the correct attributes: Verify that the response includes the attributes "name" and "eatType" as requested in the question. 2. Verify the accuracy of the attribute values: Check if the values provided for the attributes match the information given in the original text. 3. Check for any additional or missing information: Ensure that the response does not include any extra attributes or information not requested in the question. Step-by-step reference answer: 1. The response should include the attributes "name" and "eatType". 2. The values for the attributes should be: * name: The Phoenix * eatType: French pub 3. The response should not include any additional attributes or information. [End of Evaluation Plan] [Start of Plan Execution] Step 1: Check if the response provides the correct attributes Assistant A's response includes the attributes "Name", "eatType", "location", and "priceRange". However, the question only requests the attributes "nameänd "eatType". Assistant B's response includes the attributes "name" and "eatType", which matches the request. Step 2: Verify the accuracy of the attribute values ... Step 3: Check for any additional or missing information Assistant A's response includes additional attributes "location" and "priceRange" that are not requested in the question. Assistant B's response does not include any additional attributes or information.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4dc776fc-2227-4a16-b045-264f39a67143Cited by top-tier papers31
- Skywork-Reward-V2: Scaling Preference Data Curation via Human-AI SynergyChris Yuhao Liu, Liang Zeng, Yuzhen Xiao, Jujie He et al.ICLR 2026 · 211 citations
- ReMA: Learning to Meta-Think for LLMs with Multi-agent Reinforcement LearningZiyu Wan, Yunxiang Li, Xiaoyu Wen, Yan Song et al.NeurIPS 2025 · 76 citations
- J1: Incentivizing Thinking in LLM-as-a-Judge via Reinforcement LearningChenxi Whitehouse, Tianlu Wang, Ping Yu, Xian Li et al.ICLR 2026 · 74 citations
- DeepReview: Improving LLM-based Paper Review with Human-like Deep Thinking ProcessMinjun Zhu, Yixuan Weng, Linyi Yang, Yue ZhangACL 2025 · 70 citations
- ReasonFlux-PRM: Trajectory-Aware PRMs for Long Chain-of-Thought Reasoning in LLMsJiaru Zou, Ling Yang, Jingwen Gu, Jiahao Qiu et al.NeurIPS 2025 · 51 citations
Builds on26
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- STaR: Bootstrapping Reasoning With ReasoningEric Zelikman, Yuhuai Wu, Jesse Mu, Noah D. GoodmanNeurIPS 2022 · 1,126 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
Related papers
- Think-J: Learning to Think for Generative LLM-as-a-JudgeHui Huang, Yancheng He, Hongli Zhou, Rui Zhang et al.AAAI 2026 · 12 citations
- EvolvR: Self-Evolving Pairwise Reasoning for Story Evaluation to Enhance GenerationXinda Wang, Zhengxu Hou, Yangshijie Zhang, Bingren Yan et al.ACL 2026 · 1 citation
- Learning LLM-as-a-Judge for Preference AlignmentZiyi Ye, Xiangsheng Li, Qiuchi Li, Qingyao Ai et al.ICLR 2025
- Rectifying LLM Thought from Lens of OptimizationJunnan Liu, Hongwei Liu, Songyang Zhang, Kai ChenICLR 2026 · 3 citations
- Thinking LLMs: General Instruction Following with Thought GenerationTianhao Wu, Janice Lan, Weizhe Yuan, Jiantao Jiao et al.ICML 2025 · 3 citations
