XLQA: A Benchmark for Locale-Aware Multilingual Open-Domain Question Answering
Keon-Woo Roh, Yeong-Joon Ju, Seong-Whan Lee
Abstract
Large Language Models (LLMs) have shown significant progress in Open-Domain Question Answering (ODQA), yet most evaluations focus on English and assume locale-invariant answers across languages.This assumption neglects the cultural and regional variations that affect question understanding and answer, leading to biased evaluation in multilingual benchmarks.To address these limitations, we introduce XLQA, a novel benchmark explicitly designed for locale-sensitive multilingual ODQA.XLQA contains 3,000 English seed questions expanded to eight languages, with careful filtering for semantic consistency and human-verified annotations distinguishing locale-invariant and locale-sensitive cases.Our evaluation of five state-of-the-art multilingual LLMs reveals notable failures on localesensitive questions, exposing gaps between English and other languages due to a lack of locale-grounding knowledge.We provide a systematic framework and scalable methodology for assessing multilingual QA under diverse cultural contexts, offering a critical resource to advance the real-world applicability of multilingual ODQA systems.Our findings suggest that disparities in training data distribution contribute to differences in both linguistic competence and locale-awareness across models.https://github.com/ro-ko/XLQA
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ef03f403-bd00-4137-b8c4-d6526bca23d3Cited by top-tier papers1
Ask how each one uses itBuilds on9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning OptimizationYidong Wang, Zhuohao Yu, Wenjin Yao, Zhengran Zeng et al.ICLR 2024 · 368 citations
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis et al.EMNLP 2020 · 142 citations
- Don't Trust ChatGPT when your Question is not in English: A Study of Multilingual Abilities and Types of LLMsXiang Zhang, Senyu Li, Bradley Hauer, Ning Shi et al.EMNLP 2023 · 58 citations
- On the Cross-lingual Transferability of Monolingual RepresentationsMikel Artetxe, Sebastian Ruder, Dani YogatamaACL 2020 · 57 citations
Related papers
- Do You Know About My Nation? Investigating Multilingual Language Models' Cultural Literacy Through Factual KnowledgeEshaan Tanwar, Anwoy Chatterjee, Michael Saxon, Alon Albalak et al.EMNLP 2025 · 4 citations
- Location Not Found: Exposing Implicit Local and Global Biases in Multilingual LLMsGuy Mor-Lan, Omer Goldman, Matan Eyal, Adi Mayrav Gilady et al.ACL 2026 · 2 citations
- MLQA: Evaluating Cross-lingual Extractive Question AnsweringPatrick Lewis, Barlas Oguz, Ruty Rinott, Sebastian Riedel et al.ACL 2020 · 52 citations
- Afri-MCQA: Multimodal Cultural Question Answering for African LanguagesAtnafu Lambebo Tonja, Srija Anand, Emilio Villa-Cueva, Israel Abebe Azime et al.ACL 2026 · 2 citations
- Truth Knows No Language: Evaluating Truthfulness Beyond EnglishBlanca Calvo Figueras, Eneko Sagarzazu, Julen Etxaniz, Jeremy Barnes et al.ACL 2025
