CaLMQA: Exploring culturally specific long-form question answering across 23 languages
Shane Arora, Marzena Karpinska, Hung-Ting Chen, Ipsita Bhattacharjee, Mohit Iyyer, Eunsol Choi
摘要
Despite rising global usage of large language models (LLMs), their ability to generate longform answers to culturally specific questions remains unexplored in many languages. To fill this gap, we perform the first study of textual multilingual long-form QA by creating CALMQA, a dataset of 51.7K culturally specific questions across 23 different languages. We define culturally specific questions as those that refer to concepts unique to one or a few cultures, or have different answers depending on the cultural or regional context. We obtain these questions by crawling naturallyoccurring questions from community web forums in high-resource languages, and by hiring native speakers to write questions in underresourced, rarely-studied languages such as Fijian and Kirundi. Our data collection methodologies are translation-free, enabling the collection of culturally unique questions like 'Kuber iki umwami wa mbere w'uburundi yitwa Ntare?" (Kirundi; English translation: "Why was the first king of Burundi called Ntare (Lion)?"). We evaluate factuality, relevance and surface-level quality of LLM-generated long-form answers, finding that (1) for many languages, even the best models make critical surface-level errors (e.g., answering in the wrong language, repetition), especially for lowresource languages; and (2) answers to culturally specific questions contain more factual errors than answers to culturally agnostic questions -questions that have consistent meaning and answer across many cultures. We release CALMQA to facilitate future research in cultural and multilingual long-form QA. github.com/2015aroras/CaLMQA hf.co/datasets/shanearora/CaLMQA © CC BY 4.0
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Do You Know About My Nation? Investigating Multilingual Language Models' Cultural Literacy Through Factual KnowledgeEshaan Tanwar, Anwoy Chatterjee, Michael Saxon, Alon Albalak 等EMNLP 2025 · 被引用 4 次
- Culture In a Frame: C3B as a Comic-Based Benchmark for Multimodal Culturally AwarenessYuchen Song, Andong Chen, Wenxin Zhu, Kehai Chen 等ICLR 2026 · 被引用 3 次
- BhashaSutra: A Task-Centric Unified Survey of Indian NLP Datasets, Corpora, and ResourcesRaghvendra Kumar, Devankar Raj, Sriparna SahaACL 2026
- CulturalBench: A Robust, Diverse and Challenging Benchmark for Measuring LMs' Cultural Knowledge Through Human-AI Red-TeamingYu Ying Chiu, Liwei Jiang, Bill Yuchen Lin, Chan Young Park 等ACL 2025
- What are Foundation Models Cooking in the Post-Soviet World?Anton Lavrouk, Tarek Naous, Alan Ritter, Wei XuEMNLP 2025
它引用的顶会 Paper11
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- MM-Vet: Evaluating Large Multimodal Models for Integrated CapabilitiesWeihao Yu, Zhengyuan Yang, Linjie Li, Jianfeng Wang 等ICML 2024 · 被引用 1,191 次
- FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text GenerationSewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis 等EMNLP 2023 · 被引用 225 次
- Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual EvaluationShivalika Singh, Angelika Romanou, Clémentine Fourrier, David Ifeoluwa Adelani 等ACL 2025 · 被引用 144 次
相关 Paper
- Can LLM Generate Culturally Relevant Commonsense QA Data? Case Study in Indonesian and SundaneseRifki Afina Putri, Faiz Ghifari Haznitrama, Dea Adhista, Alice OhEMNLP 2024 · 被引用 7 次
- XLQA: A Benchmark for Locale-Aware Multilingual Open-Domain Question AnsweringKeon-Woo Roh, Yeong-Joon Ju, Seong-Whan LeeEMNLP 2025
- Afri-MCQA: Multimodal Cultural Question Answering for African LanguagesAtnafu Lambebo Tonja, Srija Anand, Emilio Villa-Cueva, Israel Abebe Azime 等ACL 2026 · 被引用 2 次
- CulFiT: A Fine-grained Cultural-aware LLM Training Paradigm via Multilingual Critique Data SynthesisRuixiang Feng, Shen Gao, Xiuying Chen, Lisi Chen 等ACL 2025
- How Much Do LLMs Hallucinate across Languages? On Realistic Multilingual Estimation of LLM HallucinationSaad Obaid ul Islam, Anne Lauscher, Goran GlavasEMNLP 2025
