Direct Multi-Turn Preference Optimization for Language Agents
Wentao Shi, Mengqi Yuan, Junkang Wu, Qifan Wang, Fuli Feng
Abstract
Adapting Large Language Models (LLMs) for agent tasks is critical in developing language agents. Direct Preference Optimization (DPO) is a promising technique for this adaptation with the alleviation of compounding errors, offering a means to directly optimize Reinforcement Learning (RL) objectives. However, applying DPO to multi-turn tasks presents challenges due to the inability to cancel the partition function. Overcoming this obstacle involves making the partition function independent of the current state and addressing length disparities between preferred and dis-preferred trajectories. In this light, we replace the policy constraint with the state-action occupancy measure constraint in the RL objective and add length normalization to the Bradley-Terry model, yielding a novel loss function named DMPO for multi-turn agent tasks with theoretical explanations. Extensive experiments on three multi-turn agent task datasets confirm the effectiveness and superiority of the DMPO loss. The code is available at https: //github.com/swt-user/DMPO .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext af22e1c9-0d74-4332-bfa0-1168c84f50f2Cited by top-tier papers16
- Fine-Grained Preference Optimization Improves Spatial Reasoning in VLMsYifan Shen, Yuanzhe Liu, Jingyuan Zhu, Xu Cao et al.NeurIPS 2025 · 41 citations
- The Lighthouse of Language: Enhancing LLM Agents via Critique-Guided ImprovementRuihan Yang, Fanghua Ye, Jian Li, Siyu Yuan et al.NeurIPS 2025 · 21 citations
- Iterative Tool Usage Exploration for Multimodal Agents via Step-wise Preference TuningPengxiang Li, Zhi Gao, Bofei Zhang, Yapeng Mi et al.NeurIPS 2025 · 21 citations
- Benchmarking LLM Tool-Use in the WildPeijie Yu, Wei Liu, Yifan Yang, Jinjian Li et al.ICLR 2026 · 20 citations
- ATPO: Adaptive Tree Policy Optimization for Multi-Turn Medical DialogueRuike Cao, Shaojie Bai, Fugen Yao, Liang Dong et al.ICLR 2026 · 8 citations
Builds on10
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
- Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied AgentsWenlong Huang, Pieter Abbeel, Deepak Pathak, Igor MordatchICML 2022 · 1,539 citations
- SimPO: Simple Preference Optimization with a Reference-Free RewardYu Meng, Mengzhou Xia, Danqi ChenNeurIPS 2024 · 1,203 citations
- ALFWorld: Aligning Text and Embodied Environments for Interactive LearningMohit Shridhar, Xingdi Yuan, Marc-Alexandre Côté, Yonatan Bisk et al.ICLR 2021 · 819 citations
Related papers
- Autoregressive Direct Preference OptimizationMasanari Oi, Mahiro Ukai, Masahiro Kaneko, Naoaki Okazaki et al.ICML 2026 · 1 citation
- Right Now, Wrong Then: Non-Stationary Direct Preference Optimization under Preference DriftSeongho Son, William Bankes, Sayak Ray Chowdhury, Brooks Paige et al.ICML 2025
- TokenRatio: Principled Token-Level Preference Optimization via Ratio MatchingTruong Nguyen, Tien-Phat Nguyen, Linh Van, Duy Nguyen et al.ICML 2026
- Token-level Direct Preference OptimizationYongcheng Zeng, Guoqing Liu, Weiyu Ma, Ning Yang et al.ICML 2024 · 136 citations
- Pruning Long Chain-of-Thought of Large Reasoning Models via Small-Scale Preference OptimizationBin Hong, Jiayu Liu, Kai Zhang, Jianwen Sun et al.ICLR 2026 · 1 citation
