Adversarial Language Games for Advanced Natural Language Intelligence
Yuan Yao, Haoxi Zhong, Zhengyan Zhang, Xu Han, Xiaozhi Wang, Kai Zhang, Chaojun Xiao, Guoyang Zeng, Zhiyuan Liu, Maosong Sun
摘要
We study the problem of adversarial language games 1 , in which multiple agents with conflicting goals compete with each other via natural language interactions. While adversarial language games are ubiquitous in human activities, little attention has been devoted to this field in natural language processing. In this work, we propose a challenging adversarial language game called Adversarial Taboo as an example, in which an attacker and a defender compete around a target word. The attacker is tasked with inducing the defender to utter the target word invisible to the defender, while the defender is tasked with detecting the target word before being induced by the attacker. In Adversarial Taboo, a successful attacker and defender need to hide or infer the intention, and induce or defend during conversations. This requires several advanced language abilities, such as adversarial pragmatic reasoning and goal-oriented language interactions in open domain, which will facilitate many downstream NLP tasks. To instantiate the game, we create a game environment and a competition platform. Comprehensive experiments on several baseline attack and defense strategies show promising and interesting results, based on which we discuss some directions for future research. The code and datasets of this paper can be obtained from https://github.com/thunlp/AdversarialTaboo .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Self-playing Adversarial Language Game Enhances LLM ReasoningPengyu Cheng, Tianhao Hu, Han Xu, Zhisong Zhang 等NeurIPS 2024 · 被引用 120 次
- CoMet: Metaphor-Driven Covert Communication for Multi-Agent Language GamesShuhang Xu, Fangwei ZhongACL 2025 · 被引用 2 次
- Foresight Optimization for Strategic Reasoning in Large Language ModelsJessie Wang, Jiawen Duan, Jian Wang, Kaitao Song 等ACL 2026
它引用的顶会 Paper1
相关 Paper
- Inverse Reinforcement Learning with Natural Language GoalsLi Zhou, Kevin SmallAAAI 2021 · 被引用 40 次
- Safety Alignment of LMs via Non-cooperative GamesAnselm Paulus, Ilia Kulikov, Brandon Amos, REMI MUNOS 等ICML 2026 · 被引用 4 次
- Why Should Adversarial Perturbations be Imperceptible? Rethink the Research Paradigm in Adversarial NLPYangyi Chen, Hongcheng Gao, Ganqu Cui, Fanchao Qi 等EMNLP 2022 · 被引用 28 次
- Explaining Decisions of Agents in Mixed-Motive GamesMaayan Orner, Oleg Maksimov, Akiva Kleinerman, Charles Ortiz 等AAAI 2025 · 被引用 4 次
- Adversarial Regularization as Stackelberg Game: An Unrolled Optimization ApproachSimiao Zuo, Chen Liang, Haoming Jiang, Xiaodong Liu 等EMNLP 2021 · 被引用 4 次
