Battling against Tough Resister: Strategy Planning with Adversarial Game for Non-collaborative Dialogues
Haiyang Wang, Zhiliang Tian, Yuchen Pan, Xin Song, Xin Niu, Minlie Huang, Bin Zhou
Abstract
Non-collaborative dialogue involves two participants with conflicting interests engaging in a multi-round dialogue to achieve their own goals. Strategy planning is the key to guiding both participants towards a consensus. Most LLMs-based methods use stimulus prompts or external strategy planners for strategy planning. However, stimulus prompts fail to teach LLMs to plan dialogue strategies explicitly. Moreover, training external strategy planners doesn’t fully account for adversarial interactions, thereby limiting their effectiveness against tough resisters. In this paper, to mitigate the above issues, we propose GAIA , a G ame-based A dversarial self-play I nter A ctive training paradigm, which constructs an adversarial two-player (a persuader and a resister) zero-sum game and guides the game to approximate Nash Equilibrium (NE) via reinforcement learning (RL) for the non-collaborative dialogues. First, we design a Chain-of-Mind prompt to reason the resister’s dialogue act step-by-step to plan the persuasive strategies. Secondly, to adversarially improve the persuader, we construct diverse resistant planners and theoretically improve the persuader’s optimal lower bound. Finally, we iteratively optimise their policies via adversarial self-play interactive RL and design an ϵ -NE verification algorithm to approximate the game’s NE. Experiments on three datasets show that our model obtains state-of-the-art performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on11
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Evaluating and Inducing Personality in Pre-trained Language ModelsGuangyuan Jiang, Manjie Xu, Song-Chun Zhu, Wenjuan Han et al.NeurIPS 2023 · 192 citations
- Plug-and-Play Policy Planner for Large Language Model Powered Dialogue AgentsYang Deng, Wenxuan Zhang, Wai Lam, See-Kiong Ng et al.ICLR 2024 · 86 citations
- DialoGraph: Incorporating Interpretable Strategy-Graph Networks into Negotiation DialoguesRishabh Joshi, Vidhisha Balachandran, Shikhar Vashishth, Alan W. Black et al.ICLR 2021 · 39 citations
- Augmenting Non-Collaborative Dialog Systems with Explicit Semantic and Strategic Dialog HistoryYiheng Zhou, Yulia Tsvetkov, Alan W. Black, Zhou YuICLR 2020 · 35 citations
Related papers
- AGD: Adversarial Game Defense Against Jailbreak Attacks in Large Language ModelsShilong Pan, Zhiliang Tian, Zhen Huang, Wanlong Yu et al.ACL 2025 · 2 citations
- Self-playing Adversarial Language Game Enhances LLM ReasoningPengyu Cheng, Tianhao Hu, Han Xu, Zhisong Zhang et al.NeurIPS 2024 · 120 citations
- Reward-Guided Prompt Evolving in Reinforcement Learning for LLMsZiyu Ye, Rishabh Agarwal, Tianqi Liu, Rishabh Joshi et al.ICML 2025
- D²Plan: Dual-Agent Dynamic Global Planning for Complex Retrieval-Augmented ReasoningKangcheng Luo, Tinglang Wu, Yansong FengACL 2026
- Strategist: Self-improvement of LLM Decision Making via Bi-Level Tree SearchJonathan Light, Min Cai, Weiqin Chen, Guanzhi Wang et al.ICLR 2025
