Battling against Tough Resister: Strategy Planning with Adversarial Game for Non-collaborative Dialogues
Haiyang Wang, Zhiliang Tian, Yuchen Pan, Xin Song, Xin Niu, Minlie Huang, Bin Zhou
摘要
Non-collaborative dialogue involves two participants with conflicting interests engaging in a multi-round dialogue to achieve their own goals. Strategy planning is the key to guiding both participants towards a consensus. Most LLMs-based methods use stimulus prompts or external strategy planners for strategy planning. However, stimulus prompts fail to teach LLMs to plan dialogue strategies explicitly. Moreover, training external strategy planners doesn’t fully account for adversarial interactions, thereby limiting their effectiveness against tough resisters. In this paper, to mitigate the above issues, we propose GAIA , a G ame-based A dversarial self-play I nter A ctive training paradigm, which constructs an adversarial two-player (a persuader and a resister) zero-sum game and guides the game to approximate Nash Equilibrium (NE) via reinforcement learning (RL) for the non-collaborative dialogues. First, we design a Chain-of-Mind prompt to reason the resister’s dialogue act step-by-step to plan the persuasive strategies. Secondly, to adversarially improve the persuader, we construct diverse resistant planners and theoretically improve the persuader’s optimal lower bound. Finally, we iteratively optimise their policies via adversarial self-play interactive RL and design an ϵ -NE verification algorithm to approximate the game’s NE. Experiments on three datasets show that our model obtains state-of-the-art performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper11
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Evaluating and Inducing Personality in Pre-trained Language ModelsGuangyuan Jiang, Manjie Xu, Song-Chun Zhu, Wenjuan Han 等NeurIPS 2023 · 被引用 192 次
- Plug-and-Play Policy Planner for Large Language Model Powered Dialogue AgentsYang Deng, Wenxuan Zhang, Wai Lam, See-Kiong Ng 等ICLR 2024 · 被引用 86 次
- DialoGraph: Incorporating Interpretable Strategy-Graph Networks into Negotiation DialoguesRishabh Joshi, Vidhisha Balachandran, Shikhar Vashishth, Alan W. Black 等ICLR 2021 · 被引用 39 次
- Augmenting Non-Collaborative Dialog Systems with Explicit Semantic and Strategic Dialog HistoryYiheng Zhou, Yulia Tsvetkov, Alan W. Black, Zhou YuICLR 2020 · 被引用 35 次
相关 Paper
- AGD: Adversarial Game Defense Against Jailbreak Attacks in Large Language ModelsShilong Pan, Zhiliang Tian, Zhen Huang, Wanlong Yu 等ACL 2025 · 被引用 2 次
- Self-playing Adversarial Language Game Enhances LLM ReasoningPengyu Cheng, Tianhao Hu, Han Xu, Zhisong Zhang 等NeurIPS 2024 · 被引用 120 次
- Reward-Guided Prompt Evolving in Reinforcement Learning for LLMsZiyu Ye, Rishabh Agarwal, Tianqi Liu, Rishabh Joshi 等ICML 2025
- D²Plan: Dual-Agent Dynamic Global Planning for Complex Retrieval-Augmented ReasoningKangcheng Luo, Tinglang Wu, Yansong FengACL 2026
- Strategist: Self-improvement of LLM Decision Making via Bi-Level Tree SearchJonathan Light, Min Cai, Weiqin Chen, Guanzhi Wang 等ICLR 2025
