SOTOPIA: Interactive Evaluation for Social Intelligence in Language Agents
Xuhui Zhou, Hao Zhu, Leena Mathur, Ruohong Zhang, Haofei Yu, Zhengyang Qi, Louis-Philippe Morency, Yonatan Bisk, Daniel Fried, Graham Neubig, Maarten Sap
Abstract
Humans are social beings; we pursue social goals in our daily interactions, which is a crucial aspect of social intelligence. Yet, AI systems' abilities in this realm remain elusive. We present SOTOPIA, an open-ended environment to simulate complex social interactions between artificial agents and evaluate their social intelligence. In our environment, agents role-play and interact under a wide variety of scenarios; they coordinate, collaborate, exchange, and compete with each other to achieve complex social goals. We simulate the role-play interaction between LLM-based agents and humans within this task space and evaluate their performance with a holistic evaluation framework called SOTOPIA-Eval. With SOTOPIA, we find significant differences between these models in terms of their social intelligence, and we identify a subset of SOTOPIA scenarios, SOTOPIA-hard, that is generally challenging for all models. We find that on this subset, GPT-4 achieves a significantly lower goal completion rate than humans and struggles to exhibit social commonsense reasoning and strategic communication skills. These findings demonstrate SOTOPIA's promise as a general platform for research on evaluating and improving social intelligence in artificial agents.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers50
- Can Large Language Model Agents Simulate Human Trust Behavior?Chengxing Xie, Canyu Chen, Feiran Jia, Ziyu Ye et al.NeurIPS 2024 · 183 citations
- OpenAgentSafety: A Comprehensive Framework For Evaluating Real-World AI Agent SafetySanidhya Vijayvargiya, Aditya Bharat Soni, Xuhui Zhou, Zora Zhiruo Wang et al.ICLR 2026 · 75 citations
- Simulacrum of Stories: Examining Large Language Models as Qualitative Research ParticipantsShivani Kapania, William Agnew, Motahhare Eslami, Hoda Heidari et al.CHI 2025 · 59 citations
- Evaluating Language Model Agency Through NegotiationsTim R. Davidson, Veniamin Veselovsky, Michal Kosinski, Robert WestICLR 2024 · 52 citations
- Self-Alignment of Large Language Models via Monopolylogue-based Social Scene SimulationXianghe Pang, Shuo Tang, Rui Ye, Yuxin Xiong et al.ICML 2024 · 50 citations
Builds on24
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model SocietyGuohao Li, Hasan Hammoud, Hani Itani, Dmitrii Khizbullin et al.NeurIPS 2023 · 1,975 citations
- ChatEval: Towards Better LLM-based Evaluators through Multi-Agent DebateChi-Min Chan, Weize Chen, Yusheng Su, Jianxuan Yu et al.ICLR 2024 · 871 citations
- A Simple Language Model for Task-Oriented DialogueEhsan Hosseini-Asl, Bryan McCann, Chien-Sheng Wu, Semih Yavuz et al.NeurIPS 2020 · 590 citations
Related papers
- SOTOPIA-π: Interactive Learning of Socially Intelligent Language AgentsRuiyi Wang, Haofei Yu, Wenxin Zhang, Zhengyang Qi et al.ACL 2024
- SDPO: Segment-Level Direct Preference Optimization for Social AgentsAobo Kong, Wentao Ma, Shiwan Zhao, Yongbin Li et al.ACL 2025
- Spontaneous Giving and Calculated Greed in Language ModelsYuxuan Li, Hirokazu ShiradoEMNLP 2025 · 1 citation
- Neural Theory-of-Mind? On the Limits of Social Intelligence in Large LMsMaarten Sap, Ronan Le Bras, Daniel Fried, Yejin ChoiEMNLP 2022 · 92 citations
- ALSO: Adversarial Online Strategy Optimization for Social AgentsXiang Li, Liping Yi, Mingze Kong, Min Zhang et al.ICML 2026
