Susu Box or Piggy Bank: Assessing Cultural Commonsense Knowledge between Ghana and the US
Christabel Acquaye, Haozhe An, Rachel Rudinger
Abstract
Recent work has highlighted the culturallycontingent nature of commonsense knowledge (Shen et al., 2024) . We introduce AMAMMERE (/A:.mA:.mu:.reI/), from the Akan word meaning 'culture.' This test set of 525 multiple-choice questions is designed to evaluate the commonsense knowledge of English LLMs, relative to the cultural contexts of Ghana and the United States. To create AMAMMERE, we select a set of multiplechoice questions (MCQs) from existing commonsense datasets and rewrite them in a multistage process involving surveys of Ghanaian and U.S. participants. In three rounds of surveys, participants from both pools are solicited to (1) write correct and incorrect answer choices, ( 2 ) rate individual answer choices on a 5-point Likert scale, and (3) select the best answer choice from the newly-constructed MCQ items, in a final validation step. By engaging participants at multiple stages, our procedure ensures that participant perspectives are incorporated both in the creation and validation of test items, resulting in high levels of agreement within each pool. We evaluate several off-the-shelf English LLMs on AMAMMERE. 1 Uniformly, models prefer answers choices that align with the preferences of U.S. annotators over Ghanaian annotators. Additionally, when test items specify a cultural context (Ghana or the U.S.), models exhibit some ability to adapt, but performance is consistently better in U.S. contexts than Ghanaian. As large resources are devoted to the advancement of English LLMs, our findings underscore the need for culturally adaptable models and evaluations to meet the needs of diverse English-speaking populations around the world. Original Version Ghanaian Version Answer Choices Unspecified Version American Version Goal: How do you dry herbs when making homemade herb oil? sol1: Place freshly washed and clean herbs in a cold refrigerator sol2: Place freshly washed and clean herbs in a food dehydrator Context: This person plans to make homemade herb oil. Queson: How will you typically dry the herbs for this? Context: This person plans to make homemade herb oil. Queson: How will Kpakpo typically dry the herbs for this in Ghana? Context: This person plans to make homemade herb oil. Queson: How will Zach typically dry the herbs for this in the USA? A. Place the herbs in a basket for the sun to dry it B. Place the herbs in a food dehydrator to dry it C. Put the herbs in an electric dryer to dry it D. Dunk the herbs in some water to dry it
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 432375c1-249a-48e8-97fa-d9090b786ad0Cited by top-tier papers4
- Common to Whom? Regional Cultural Commonsense and LLM Bias in IndiaSangmitra Madhusudan, Trush Shashank More, Steph Buongiorno, Renata Dividino et al.ACL 2026 · 1 citation
- On the Mutual Influence of Gender and Occupation in LLM RepresentationsHaozhe An, Connor Baumler, Abhilasha Sancheti, Rachel RudingerACL 2025
- Which of These Best Describes Multiple Choice Evaluation with LLMs? A) Forced B) Flawed C) Fixable D) All of the AboveNishant Balepur, Rachel Rudinger, Jordan Lee Boyd-GraberACL 2025
- Arguments that Alter Minds: LLM Rationales Sway Human (and LLM) Notions of PlausibilityShramay Palta, Peter Rankel, Sarah Wiegreffe, Rachel RudingerACL 2026
Builds on6
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- Broaden the Vision: Geo-Diverse Visual Commonsense ReasoningDa Yin, Liunian Harold Li, Ziniu Hu, Nanyun Peng et al.EMNLP 2021 · 32 citations
- XCOPA: A Multilingual Dataset for Causal Commonsense ReasoningEdoardo Maria Ponti, Goran Glavas, Olga Majewska, Qianchu Liu et al.EMNLP 2020 · 6 citations
- Artifacts or Abduction: How Do LLMs Answer Multiple-Choice Questions Without the Question?Nishant Balepur, Abhilasha Ravichander, Rachel RudingerACL 2024 · 5 citations
- Having Beer after Prayer? Measuring Cultural Bias in Large Language ModelsTarek Naous, Michael J. Ryan, Alan Ritter, Wei XuACL 2024
Related papers
- CulturalBench: A Robust, Diverse and Challenging Benchmark for Measuring LMs' Cultural Knowledge Through Human-AI Red-TeamingYu Ying Chiu, Liwei Jiang, Bill Yuchen Lin, Chan Young Park et al.ACL 2025
- GeoMLAMA: Geo-Diverse Commonsense Probing on Multilingual Pre-Trained Language ModelsDa Yin, Hritik Bansal, Masoud Monajatipoor, Liunian Harold Li et al.EMNLP 2022 · 27 citations
- Africa Health Check: Probing Cultural Bias in Medical LLMsCharles Nimo, Shuheng Liu, Irfan Essa, Michael L. BestEMNLP 2025
- MMLU-CF: A Contamination-free Multi-task Language Understanding BenchmarkQihao Zhao, Yangyu Huang, Tengchao Lv, Lei Cui et al.ACL 2025 · 33 citations
- Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual EvaluationShivalika Singh, Angelika Romanou, Clémentine Fourrier, David Ifeoluwa Adelani et al.ACL 2025 · 144 citations
