Adversarial Language Games for Advanced Natural Language Intelligence
Yuan Yao, Haoxi Zhong, Zhengyan Zhang, Xu Han, Xiaozhi Wang, Kai Zhang, Chaojun Xiao, Guoyang Zeng, Zhiyuan Liu, Maosong Sun
Abstract
We study the problem of adversarial language games 1 , in which multiple agents with conflicting goals compete with each other via natural language interactions. While adversarial language games are ubiquitous in human activities, little attention has been devoted to this field in natural language processing. In this work, we propose a challenging adversarial language game called Adversarial Taboo as an example, in which an attacker and a defender compete around a target word. The attacker is tasked with inducing the defender to utter the target word invisible to the defender, while the defender is tasked with detecting the target word before being induced by the attacker. In Adversarial Taboo, a successful attacker and defender need to hide or infer the intention, and induce or defend during conversations. This requires several advanced language abilities, such as adversarial pragmatic reasoning and goal-oriented language interactions in open domain, which will facilitate many downstream NLP tasks. To instantiate the game, we create a game environment and a competition platform. Comprehensive experiments on several baseline attack and defense strategies show promising and interesting results, based on which we discuss some directions for future research. The code and datasets of this paper can be obtained from https://github.com/thunlp/AdversarialTaboo .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Self-playing Adversarial Language Game Enhances LLM ReasoningPengyu Cheng, Tianhao Hu, Han Xu, Zhisong Zhang et al.NeurIPS 2024 · 120 citations
- CoMet: Metaphor-Driven Covert Communication for Multi-Agent Language GamesShuhang Xu, Fangwei ZhongACL 2025 · 2 citations
- Foresight Optimization for Strategic Reasoning in Large Language ModelsJessie Wang, Jiawen Duan, Jian Wang, Kaitao Song et al.ACL 2026
Builds on1
Related papers
- Inverse Reinforcement Learning with Natural Language GoalsLi Zhou, Kevin SmallAAAI 2021 · 40 citations
- Safety Alignment of LMs via Non-cooperative GamesAnselm Paulus, Ilia Kulikov, Brandon Amos, REMI MUNOS et al.ICML 2026 · 4 citations
- Why Should Adversarial Perturbations be Imperceptible? Rethink the Research Paradigm in Adversarial NLPYangyi Chen, Hongcheng Gao, Ganqu Cui, Fanchao Qi et al.EMNLP 2022 · 28 citations
- Explaining Decisions of Agents in Mixed-Motive GamesMaayan Orner, Oleg Maksimov, Akiva Kleinerman, Charles Ortiz et al.AAAI 2025 · 4 citations
- Adversarial Regularization as Stackelberg Game: An Unrolled Optimization ApproachSimiao Zuo, Chen Liang, Haoming Jiang, Xiaodong Liu et al.EMNLP 2021 · 4 citations
