SocialCC: Interactive Evaluation for Cultural Competence in Language Agents
Jincenzi Wu, Jianxun Lian, Dingdong Wang, Helen M. Meng
摘要
Large Language Models (LLMs) are increasingly deployed worldwide, yet their ability to navigate cultural nuances remains underex-plored. Misinterpreting cultural content can lead to AI-generated responses that are offensive or inappropriate, limiting their usability in global applications such as customer service, diplomatic communication, and online education. While prior research has evaluated cultural knowledge of LLMs, existing benchmarks fail to assess dynamic cultural competence — the ability to apply cultural knowledge effectively in real-world interactions. To address this gap, we introduce SocialCC , a novel benchmark designed to evaluate cultural competence through multi-turn interactive intercultural scenarios. It comprises 3,060 human-written scenarios spanning 60 countries across six continents. Through extensive experiments on eight prominent LLMs, our findings reveal a significant gap between the cultural knowledge stored in these models and their ability to apply it effectively in cross-cultural communication. We release our code and data at https: //github.com/jincenziwu/SocialCC .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper6
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- Not All Countries Celebrate Thanksgiving: On the Cultural Dominance in Large Language ModelsWenxuan Wang, Wenxiang Jiao, Jingyuan Huang, Ruyi Dai 等ACL 2024 · 被引用 21 次
- Towards Measuring and Modeling "Culture" in LLMs: A SurveyMuhammad Farid Adilazuarda, Sagnik Mukherjee, Pradhyumna Lavania, Siddhant Singh 等EMNLP 2024 · 被引用 21 次
- COKE: A Cognitive Knowledge Graph for Machine Theory of MindJincenzi Wu, Zhuang Chen, Jiawen Deng, Sahand Sabour 等ACL 2024
相关 Paper
- LiveCultureBench: a Multi-Agent, Multi-Cultural Benchmark for Large Language Models in Dynamic Social SimulationsViet Thanh Pham, Lizhen Qu, Thuy-Trang Vu, Gholamreza Haffari 等ACL 2026
- Culture In a Frame: C3B as a Comic-Based Benchmark for Multimodal Culturally AwarenessYuchen Song, Andong Chen, Wenxin Zhu, Kehai Chen 等ICLR 2026 · 被引用 3 次
- Towards Dynamic Theory of Mind: Evaluating LLM Adaptation to Temporal Evolution of Human StatesYang Xiao, Jiashuo Wang, Qiancheng Xu, Changhe Song 等ACL 2025 · 被引用 12 次
- We Politely Insist: Your LLM Must Learn the Persian Art of TaarofNikta Gohari Sadr, Sahar Heidariasl, Karine Megerdoomian, Laleh Seyyed-Kalantari 等EMNLP 2025 · 被引用 2 次
- The GaoYao Benchmark: A Comprehensive Framework for Evaluating Multilingual and Multicultural Abilities of Large Language ModelsYilun Liu, Chunguang Zhao, Mengyao Piao, Lingqi Miao 等ACL 2026
