SDPO: Segment-Level Direct Preference Optimization for Social Agents
Aobo Kong, Wentao Ma, Shiwan Zhao, Yongbin Li, Yuchuan Wu, Ke Wang, Xiaoqian Liu, Qicheng Li, Yong Qin, Fei Huang
摘要
Social agents powered by large language models (LLMs) can simulate human social behaviors but fall short in handling complex goaloriented social dialogues. Direct Preference Optimization (DPO) has proven effective in aligning LLM behavior with human preferences across a variety of agent tasks. Existing DPO-based approaches for multi-turn interactions are divided into turn-level and sessionlevel methods. The turn-level method is overly fine-grained, focusing exclusively on individual turns, while session-level methods are too coarse-grained, often introducing training noise. To address these limitations, we propose Segment-Level Direct Preference Optimization (SDPO), which focuses on specific key segments within interactions to optimize multiturn agent behavior while minimizing training noise. Evaluations on the SOTOPIA benchmark demonstrate that SDPO-tuned agents consistently outperform both existing DPO-based methods and proprietary LLMs like GPT-4o, underscoring SDPO's potential to advance the social intelligence of LLM-based agents. We release our code and data at this url.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- DenseDPO: Fine-Grained Temporal Preference Optimization for Video Diffusion ModelsZiyi Wu, Anil Kag, Ivan Skorokhodov, Willi Menapace 等NeurIPS 2025 · 被引用 36 次
- Adaptive Social Learning via Mode Policy Optimization for Language AgentsMinzheng Wang, Yongbin Li, Haobo Wang, Xinghua Zhang 等ICLR 2026 · 被引用 15 次
- General Exploratory Bonus for Optimistic Exploration in RLHFWendi Li, Changdae Oh, Sharon LiICLR 2026 · 被引用 3 次
- Spend Search Where It Pays: Value-Guided Structured Sampling and Optimization for Generative RecommendationJie Jiang, Yangru Huang, Zeyu Wang, Changping Wang 等KDD 2026 · 被引用 3 次
- TO-GATE: Clarifying Questions and Summarizing Responses with Trajectory Optimization for Eliciting Human PreferenceYulin Dou, Jiangming LiuAAAI 2026 · 被引用 1 次
它引用的顶会 Paper9
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris 等UIST 2023 · 被引用 1,882 次
- WizardLM: Empowering Large Pre-Trained Language Models to Follow Complex InstructionsCan Xu, Qingfeng Sun, Kai Zheng, Xiubo Geng 等ICLR 2024 · 被引用 1,206 次
- SOTOPIA: Interactive Evaluation for Social Intelligence in Language AgentsXuhui Zhou, Hao Zhu, Leena Mathur, Ruohong Zhang 等ICLR 2024 · 被引用 288 次
- Token-level Direct Preference OptimizationYongcheng Zeng, Guoqing Liu, Weiyu Ma, Ning Yang 等ICML 2024 · 被引用 136 次
相关 Paper
- SOTOPIA-π: Interactive Learning of Socially Intelligent Language AgentsRuiyi Wang, Haofei Yu, Wenxin Zhang, Zhengyang Qi 等ACL 2024
- ALSO: Adversarial Online Strategy Optimization for Social AgentsXiang Li, Liping Yi, Mingze Kong, Min Zhang 等ICML 2026
- Solving the Granularity Mismatch: Hierarchical Preference Learning for Long-Horizon LLM AgentsHeyang Gao, Zexu Sun, Erxue Min, Hengyi Cai 等ICLR 2026 · 被引用 5 次
- MallowsPO: Fine-Tune Your LLM with Preference DispersionsHaoxian Chen, Hanyang Zhao, Henry Lam, David D. Yao 等ICLR 2025 · 被引用 1 次
- Structured Policy Optimization: Enhance Large Vision-Language Model via Self-Referenced DialogueGuohao Sun, Can Qin, Yihao Feng, Zeyuan Chen 等ICCV 2025 · 被引用 1 次
