Towards Measuring and Modeling "Culture" in LLMs: A Survey
Muhammad Farid Adilazuarda, Sagnik Mukherjee, Pradhyumna Lavania, Siddhant Singh, Alham Fikri Aji, Jacki O'Neill, Ashutosh Modi, Monojit Choudhury
摘要
We present a survey of more than 90 recent papers that aim to study cultural representation and inclusion in large language models (LLMs). We observe that none of the studies explicitly define "culture", which is a complex, multifaceted concept; instead, they probe the models on some specially designed datasets which represent certain aspects of "culture." We call these aspects the proxies of culture, and organize them across two dimensions of demographic and semantic proxies. We also categorize the probing methods employed. Our analysis indicates that only certain aspects of "culture," such as values and objectives, have been studied, leaving several other interesting and important facets, especially the multitude of semantic domains (Thompson et al., 2020) and aboutness (Hershcovich et al., 2022) , unexplored. Two other crucial gaps are the lack of robustness of probing techniques and situated studies on the impact of cultural misand under-representation in LLM-based applications. Compilation and details of papers used for the survey can be found via our GitHub repository 1 * Equal contribution 1 https://github.com/faridlazuarda/ cultural-llm-papers
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper38
- AI Suggestions Homogenize Writing Toward Western Styles and Diminish Cultural NuancesDhruv Agarwal, Mor Naaman, Aditya VashisthaCHI 2025 · 被引用 93 次
- Cultivating Pluralism In Algorithmic Monoculture: The Community Alignment DatasetLily H Zhang, Smitha Milli, Karen Long Jusko, Jonathan Smith 等ICLR 2026 · 被引用 41 次
- Culture is Not Trivia: Sociocultural Theory for Cultural NLPNaitian Zhou, David Bamman, Isaac L. BleamanACL 2025 · 被引用 33 次
- Cultural Learning-Based Culture Adaptation of Language ModelsChen Cecilia Liu, Anna Korhonen, Iryna GurevychACL 2025 · 被引用 14 次
- CAReDiO: Enhancing Cultural Alignment of LLM via Representativeness and Distinctiveness Guided Data OptimizationJing Yao, Xiaoyuan Yi, Jindong Wang, Zhicheng Dou 等ICML 2026 · 被引用 9 次
它引用的顶会 Paper30
- Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formattingMelanie Sclar, Yejin Choi, Yulia Tsvetkov, Alane SuhrICLR 2024 · 被引用 682 次
- Using Large Language Models to Simulate Multiple Humans and Replicate Human Subject StudiesGati V. Aher, Rosa I. Arriaga, Adam Tauman KalaiICML 2023 · 被引用 651 次
- Benchmarking Large Language Models in Retrieval-Augmented GenerationJiawei Chen, Hongyu Lin, Xianpei Han, Le SunAAAI 2024 · 被引用 531 次
- From Pretraining Data to Language Models to Downstream Tasks: Tracking the Trails of Political Biases Leading to Unfair NLP ModelsShangbin Feng, Chan Young Park, Yuhan Liu, Yulia TsvetkovACL 2023 · 被引用 117 次
- Extracting Cultural Commonsense Knowledge at ScaleTuan-Phong Nguyen, Simon Razniewski, Aparna S. Varde, Gerhard WeikumWWW 2023 · 被引用 102 次
相关 Paper
- From Surveys to Narratives: Rethinking Cultural Value Adaptation in LLMsMuhammad Farid Adilazuarda, Chen Cecilia Liu, Iryna Gurevych, Alham Fikri AjiEMNLP 2025
- Investigating Cultural Alignment of Large Language ModelsBadr AlKhamissi, Muhammad N. ElNokrashy, Mai Alkhamissi, Mona T. DiabACL 2024 · 被引用 27 次
- Not All Countries Celebrate Thanksgiving: On the Cultural Dominance in Large Language ModelsWenxuan Wang, Wenxiang Jiao, Jingyuan Huang, Ruyi Dai 等ACL 2024 · 被引用 21 次
- Incorporating Diverse Perspectives in Cultural Alignment: Survey of Evaluation Benchmarks Through A Three-Dimensional FrameworkMeng-Chen Wu, Si-Chi Chin, Tess Wood, Ayush Goyal 等EMNLP 2025
- Cultural Conditioning or Placebo? On the Effectiveness of Socio-Demographic PromptingSagnik Mukherjee, Muhammad Farid Adilazuarda, Sunayana Sitaram, Kalika Bali 等EMNLP 2024 · 被引用 3 次
