SDPO: Segment-Level Direct Preference Optimization for Social Agents
Aobo Kong, Wentao Ma, Shiwan Zhao, Yongbin Li, Yuchuan Wu, Ke Wang, Xiaoqian Liu, Qicheng Li, Yong Qin, Fei Huang
Abstract
Social agents powered by large language models (LLMs) can simulate human social behaviors but fall short in handling complex goaloriented social dialogues. Direct Preference Optimization (DPO) has proven effective in aligning LLM behavior with human preferences across a variety of agent tasks. Existing DPO-based approaches for multi-turn interactions are divided into turn-level and sessionlevel methods. The turn-level method is overly fine-grained, focusing exclusively on individual turns, while session-level methods are too coarse-grained, often introducing training noise. To address these limitations, we propose Segment-Level Direct Preference Optimization (SDPO), which focuses on specific key segments within interactions to optimize multiturn agent behavior while minimizing training noise. Evaluations on the SOTOPIA benchmark demonstrate that SDPO-tuned agents consistently outperform both existing DPO-based methods and proprietary LLMs like GPT-4o, underscoring SDPO's potential to advance the social intelligence of LLM-based agents. We release our code and data at this url.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3ff8397e-b723-47dd-8114-27c8991c9b0aCited by top-tier papers8
- DenseDPO: Fine-Grained Temporal Preference Optimization for Video Diffusion ModelsZiyi Wu, Anil Kag, Ivan Skorokhodov, Willi Menapace et al.NeurIPS 2025 · 36 citations
- Adaptive Social Learning via Mode Policy Optimization for Language AgentsMinzheng Wang, Yongbin Li, Haobo Wang, Xinghua Zhang et al.ICLR 2026 · 15 citations
- General Exploratory Bonus for Optimistic Exploration in RLHFWendi Li, Changdae Oh, Sharon LiICLR 2026 · 3 citations
- Spend Search Where It Pays: Value-Guided Structured Sampling and Optimization for Generative RecommendationJie Jiang, Yangru Huang, Zeyu Wang, Changping Wang et al.KDD 2026 · 3 citations
- TO-GATE: Clarifying Questions and Summarizing Responses with Trajectory Optimization for Eliciting Human PreferenceYulin Dou, Jiangming LiuAAAI 2026 · 1 citation
Builds on9
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris et al.UIST 2023 · 1,882 citations
- WizardLM: Empowering Large Pre-Trained Language Models to Follow Complex InstructionsCan Xu, Qingfeng Sun, Kai Zheng, Xiubo Geng et al.ICLR 2024 · 1,206 citations
- SOTOPIA: Interactive Evaluation for Social Intelligence in Language AgentsXuhui Zhou, Hao Zhu, Leena Mathur, Ruohong Zhang et al.ICLR 2024 · 288 citations
- Token-level Direct Preference OptimizationYongcheng Zeng, Guoqing Liu, Weiyu Ma, Ning Yang et al.ICML 2024 · 136 citations
Related papers
- SOTOPIA-π: Interactive Learning of Socially Intelligent Language AgentsRuiyi Wang, Haofei Yu, Wenxin Zhang, Zhengyang Qi et al.ACL 2024
- ALSO: Adversarial Online Strategy Optimization for Social AgentsXiang Li, Liping Yi, Mingze Kong, Min Zhang et al.ICML 2026
- Solving the Granularity Mismatch: Hierarchical Preference Learning for Long-Horizon LLM AgentsHeyang Gao, Zexu Sun, Erxue Min, Hengyi Cai et al.ICLR 2026 · 5 citations
- MallowsPO: Fine-Tune Your LLM with Preference DispersionsHaoxian Chen, Hanyang Zhao, Henry Lam, David D. Yao et al.ICLR 2025 · 1 citation
- Structured Policy Optimization: Enhance Large Vision-Language Model via Self-Referenced DialogueGuohao Sun, Can Qin, Yihao Feng, Zeyuan Chen et al.ICCV 2025 · 1 citation
