Direct Value Optimization: Improving Chain-of-Thought Reasoning in LLMs with Refined Values
Hongbo Zhang, Han Cui, Guangsheng Bao, Linyi Yang, Jun Wang, Yue Zhang
Abstract
We introduce Direct Value Optimization (DVO), an innovative reinforcement learning framework for enhancing large language models in complex reasoning tasks. Unlike traditional methods relying on preference labels, DVO utilizes value signals at individual reasoning steps, optimizing models via a mean squared error loss. The key benefit of DVO lies in its fine-grained supervision, circumventing the need for labor-intensive human annotations. Target values within the DVO are estimated using either Monte Carlo Tree Search or an outcome value model. Our empirical analysis on both mathematical and commonsense reasoning tasks shows that DVO consistently outperforms existing offline preference optimization techniques, even with fewer training steps. These findings underscore the importance of value signals in advancing reasoning capabilities and highlight DVO as a superior methodology under scenarios lacking explicit human preference information.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-FutureYidong Wang, Xin Wang, Cunxiang Wang, Junfeng Fang et al.ICML 2026 · 3 citations
- Segment-Aligned Policy Optimization for Multi-Modal ReasoningLei Gao, Zhuoming Li, Mengxi Jia, Jiakang Yuan et al.ICML 2026 · 2 citations
- Sample Lottery: Unsupervised Discovery of Critical Instances for LLM ReasoningZhiping Xiao, Yusheng Zhao, Qixin Zhang, Jiaye Xie et al.ICLR 2026
Builds on15
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- Graph of Thoughts: Solving Elaborate Problems with Large Language ModelsMaciej Besta, Nils Blach, Ales Kubicek, Robert Gerstenberger et al.AAAI 2024 · 1,292 citations
- MAmmoTH: Building Math Generalist Models through Hybrid Instruction TuningXiang Yue, Xingwei Qu, Ge Zhang, Yao Fu et al.ICLR 2024 · 558 citations
Related papers
- DAPO : Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage-Based Policy OptimizationJiacai Liu, Chaojie Wang, Chris Yuhao Liu, Liang Zeng et al.NeurIPS 2025 · 9 citations
- TPO: Aligning Large Language Models with Multi-branch & Multi-step Preference TreesWeibin Liao, Xu Chu, Yasha WangICLR 2025
- Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM ReasoningKongcheng Zhang, Qi Yao, Shunyu Liu, Yingjie Wang et al.NeurIPS 2025 · 45 citations
- A Regret Minimization Framework on Preference Learning in Large Language ModelsSuhwan Kim, Taehyun Cho, Youngsoo Jang, Geon-Hyeong Kim et al.ICML 2026
- Beyond Step Pruning: Information Theory Based Step-level Optimization for Self-Refining Large Language ModelsJinman Zhao, Erxue Min, Hui Wu, Ziheng Li et al.AAAI 2026 · 1 citation
