SocialCC: Interactive Evaluation for Cultural Competence in Language Agents
Jincenzi Wu, Jianxun Lian, Dingdong Wang, Helen M. Meng
Abstract
Large Language Models (LLMs) are increasingly deployed worldwide, yet their ability to navigate cultural nuances remains underex-plored. Misinterpreting cultural content can lead to AI-generated responses that are offensive or inappropriate, limiting their usability in global applications such as customer service, diplomatic communication, and online education. While prior research has evaluated cultural knowledge of LLMs, existing benchmarks fail to assess dynamic cultural competence — the ability to apply cultural knowledge effectively in real-world interactions. To address this gap, we introduce SocialCC , a novel benchmark designed to evaluate cultural competence through multi-turn interactive intercultural scenarios. It comprises 3,060 human-written scenarios spanning 60 countries across six continents. Through extensive experiments on eight prominent LLMs, our findings reveal a significant gap between the cultural knowledge stored in these models and their ability to apply it effectively in cross-cultural communication. We release our code and data at https: //github.com/jincenziwu/SocialCC .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 96948fd3-00e2-4e93-b484-b33d44ae6511Cited by top-tier papers1
Ask how each one uses itBuilds on6
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- Not All Countries Celebrate Thanksgiving: On the Cultural Dominance in Large Language ModelsWenxuan Wang, Wenxiang Jiao, Jingyuan Huang, Ruyi Dai et al.ACL 2024 · 21 citations
- Towards Measuring and Modeling "Culture" in LLMs: A SurveyMuhammad Farid Adilazuarda, Sagnik Mukherjee, Pradhyumna Lavania, Siddhant Singh et al.EMNLP 2024 · 21 citations
- COKE: A Cognitive Knowledge Graph for Machine Theory of MindJincenzi Wu, Zhuang Chen, Jiawen Deng, Sahand Sabour et al.ACL 2024
Related papers
- LiveCultureBench: a Multi-Agent, Multi-Cultural Benchmark for Large Language Models in Dynamic Social SimulationsViet Thanh Pham, Lizhen Qu, Thuy-Trang Vu, Gholamreza Haffari et al.ACL 2026
- Culture In a Frame: C3B as a Comic-Based Benchmark for Multimodal Culturally AwarenessYuchen Song, Andong Chen, Wenxin Zhu, Kehai Chen et al.ICLR 2026 · 3 citations
- Towards Dynamic Theory of Mind: Evaluating LLM Adaptation to Temporal Evolution of Human StatesYang Xiao, Jiashuo Wang, Qiancheng Xu, Changhe Song et al.ACL 2025 · 12 citations
- We Politely Insist: Your LLM Must Learn the Persian Art of TaarofNikta Gohari Sadr, Sahar Heidariasl, Karine Megerdoomian, Laleh Seyyed-Kalantari et al.EMNLP 2025 · 2 citations
- The GaoYao Benchmark: A Comprehensive Framework for Evaluating Multilingual and Multicultural Abilities of Large Language ModelsYilun Liu, Chunguang Zhao, Mengyao Piao, Lingqi Miao et al.ACL 2026
