Do LLMs Understand Social Knowledge? Evaluating the Sociability of Large Language Models with SocKET Benchmark
Minje Choi, Jiaxin Pei, Sagar Kumar, Chang Shu, David Jurgens
Abstract
Large language models (LLMs) have been shown to perform well at a variety of syntactic, discourse, and reasoning tasks. While LLMs are increasingly deployed in many forms including conversational agents that interact with humans, we lack a grounded benchmark to measure how well LLMs understand social language. Here, we introduce a new theorydriven benchmark, SOCKET, that contains 58 NLP tasks testing social knowledge which we group into five categories: humor & sarcasm, offensiveness, sentiment & emotion, trustworthiness, and other social factors. In tests on the benchmark, we demonstrate that current models attain only moderate performance but reveal significant potential for task transfer among different types and categories of tasks, which were predicted from theory. Through zero-shot evaluations, we show that pretrained models already possess some innate but limited capabilities of social language understanding and training on one category of tasks can improve zero-shot testing on others. Our benchmark provides a systematic way to analyze model performance on an important dimension of language and points to clear room for improvement to build more socially-aware LLMs. The resources are released at https://github. com/minjechoi/SOCKET .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 874e00de-dc73-4e74-ac3c-bb731fc14b46Cited by top-tier papers12
- DyVal: Dynamic Evaluation of Large Language Models for Reasoning TasksKaijie Zhu, Jiaao Chen, Jindong Wang, Neil Zhenqiang Gong et al.ICLR 2024 · 92 citations
- SPRIG: Improving Large Language Model Performance by System Prompt OptimizationLechen Zhang, Tolga Ergen, Lajanugen Logeswaran, Moontae Lee et al.ICLR 2026 · 35 citations
- SOTOPIA-: Dynamic Strategy Injection Learning and Social Instruction Following Evaluation for Social AgentsWenyuan Zhang, Tianyun Liu, Mengxiao Song, Xiaodong Li et al.ACL 2025 · 17 citations
- Miko: Multimodal Intention Knowledge Distillation from Large Language Models for Social-Media Commonsense DiscoveryFeihong Lu, Weiqi Wang, Yangyifei Luo, Ziqin Zhu et al.ACM MM 2024 · 11 citations
- "A good pun is its own reword": Can Large Language Models Understand Puns?Zhijun Xu, Siyu Yuan, Lingjie Chen, Deqing YangEMNLP 2024 · 6 citations
Builds on23
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained TransformersWenhui Wang, Furu Wei, Li Dong, Hangbo Bao et al.NeurIPS 2020 · 2,727 citations
- Multitask Prompted Training Enables Zero-Shot Task GeneralizationVictor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach et al.ICLR 2022 · 1,976 citations
- DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding SharingPengcheng He, Jianfeng Gao, Weizhu ChenICLR 2023 · 394 citations
Related papers
- The Skipped Beat: A Study of Sociopragmatic Understanding in LLMs for 64 LanguagesChiyu Zhang, Khai Duy Doan, Qisheng Liao, Muhammad Abdul-MageedEMNLP 2023
- SoMe: A Realistic Benchmark for LLM-based Social Media AgentsDizhan Xue, Jing Cui, Shengsheng Qian, Chuanrui Hu et al.AAAI 2026 · 1 citation
- AwarenessBench: Assessing Cognitive Capabilities of Language ModelsXiaojian Li, Rongwu Xu, Tianyun Zhang, Yue Wang et al.ACL 2026
- SocialCC: Interactive Evaluation for Cultural Competence in Language AgentsJincenzi Wu, Jianxun Lian, Dingdong Wang, Helen M. MengACL 2025 · 8 citations
- LiveCultureBench: a Multi-Agent, Multi-Cultural Benchmark for Large Language Models in Dynamic Social SimulationsViet Thanh Pham, Lizhen Qu, Thuy-Trang Vu, Gholamreza Haffari et al.ACL 2026
