Break the Checkbox: Challenging Closed-Style Evaluations of Cultural Alignment in LLMs
Mohsinul Kabir, Ajwad Abrar, Sophia Ananiadou
Abstract
A large number of studies rely on closed-style multiple-choice surveys to evaluate cultural alignment in Large Language Models (LLMs). In this work, we challenge this constrained evaluation paradigm and explore more realistic, unconstrained approaches. Using the World Values Survey (WVS) and Hofstede Cultural Dimensions as case studies, we demonstrate that LLMs exhibit stronger cultural alignment in less constrained settings, where responses are not forced. Additionally, we show that even minor changes, such as reordering survey choices, lead to inconsistent outputs, exposing the limitations of closed-style evaluations. Our findings advocate for more robust and flexible evaluation frameworks that focus on specific cultural proxies, encouraging more nuanced and accurate assessments of cultural alignment in LLMs. Imagine you are a married male from Berlin, Germany. You are 52 years of age and completed higher education level.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a8394654-d68f-4f48-99b7-a2c24d57a663Cited by top-tier papers5
- Tears or Cheers? Benchmarking LLMs via Culturally Elicited Distinct Affective ResponsesChongyuan Dai, Yaling Shen, Zihan Gao, Jia Li et al.ACL 2026 · 4 citations
- Location Not Found: Exposing Implicit Local and Global Biases in Multilingual LLMsGuy Mor-Lan, Omer Goldman, Matan Eyal, Adi Mayrav Gilady et al.ACL 2026 · 2 citations
- Distributional Open-Ended Evaluation of LLM Cultural Value Alignment Based on Value CodebookJaehyeok Lee, Xiaoyuan Yi, Jing Yao, Hyunjin Hwang et al.ICML 2026 · 1 citation
- A Game-Theoretica Negotiation Framework for Cross-Cultural ConsensusGuoxi Zhang, Jiawei Chen, Tianzhuo Yang, Jiaming Ji et al.ACL 2026
- Quantifying the Salience of Geo-Cultural Values for Pluralistic Safety AlignmentArkadiy Saakyan, Charvi Rastogi, Lora AroyoICML 2026
Builds on9
- Knowledge of cultural moral norms in large language modelsAida Ramezani, Yang XuACL 2023 · 44 citations
- The State and Fate of Linguistic Diversity and Inclusion in the NLP WorldPratik Joshi, Sebastin Santy, Amar Budhiraja, Kalika Bali et al.ACL 2020 · 40 citations
- CulturePark: Boosting Cross-cultural Understanding in Large Language ModelsCheng Li, Damien Teney, Linyi Yang, Qingsong Wen et al.NeurIPS 2024 · 34 citations
- Investigating Cultural Alignment of Large Language ModelsBadr AlKhamissi, Muhammad N. ElNokrashy, Mai Alkhamissi, Mona T. DiabACL 2024 · 27 citations
- Not All Countries Celebrate Thanksgiving: On the Cultural Dominance in Large Language ModelsWenxuan Wang, Wenxiang Jiao, Jingyuan Huang, Ruyi Dai et al.ACL 2024 · 21 citations
Related papers
- Political Compass or Spinning Arrow? Towards More Meaningful Evaluations for Values and Opinions in Large Language ModelsPaul Röttger, Valentin Hofmann, Valentina Pyatkin, Musashi Hinck et al.ACL 2024
- From Surveys to Narratives: Rethinking Cultural Value Adaptation in LLMsMuhammad Farid Adilazuarda, Chen Cecilia Liu, Iryna Gurevych, Alham Fikri AjiEMNLP 2025
- Toward Culturally Aligned LLMs through Ontology-Guided Multi-Agent ReasoningWonduk Seo, Wonseok Choi, Junseo Koh, Juhyeon Lee et al.ICML 2026 · 2 citations
- Towards Measuring and Modeling "Culture" in LLMs: A SurveyMuhammad Farid Adilazuarda, Sagnik Mukherjee, Pradhyumna Lavania, Siddhant Singh et al.EMNLP 2024 · 21 citations
- Revisiting LLM Value Probing Strategies: Are They Robust and Expressive?Siqi Shen, Mehar Singh, Lajanugen Logeswaran, Moontae Lee et al.EMNLP 2025
