Stereotype Bias in a Bilingual Setting: A Culturally Grounded Evaluation in Kazakhstan
Nurkhan Laiyk, Daniil Orel, Ayana Mussabayeva, Maiya Goloburda, Kamila Kuishibekova, Liya Goloburda, Diana Turmakhan, Preslav Nakov, Yuxia Wang, Fajri Koto
Abstract
Stereotype bias in language models has been widely examined in English, but remains largely understudied in bilingual contexts where multiple linguistic and cultural systems interact. This gap is especially important in regions where language use reflects complex historical and sociopolitical influences. In this work, we focus on Kazakhstan, a bilingual society where Kazakh, a low-resource Turkic language, and Russian, a high-resource Slavic language, are both actively used and frequently code-switched in everyday communication. We introduce Aqbileq 1 , a high-quality, humanverified dataset consisting of 5,634 stereotypebearing statements in Kazakh, Russian, and code-switched forms, covering six culturally salient domains. We evaluate both multilingual and Kazakh-specific language models using perplexity-based scoring and pretraining simulations, and find that stereotype bias is most pronounced in code-switched inputs. Our results highlight the limitations of existing evaluation frameworks and emphasize the need for culturally grounded, linguistically inclusive benchmarks to better assess and mitigate bias in language models. Warning: this paper contains example data that may be offensive, harmful, or biased.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on9
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- Masked Language Model ScoringJulian Salazar, Davis Liang, Toan Q. Nguyen, Katrin KirchhoffACL 2020 · 167 citations
- The Moral Integrity Corpus: A Benchmark for Ethical Dialogue SystemsCaleb Ziems, Jane A. Yu, Yi-Chia Wang, Alon Y. Halevy et al.ACL 2022 · 127 citations
- Code-Switching Red-Teaming: LLM Evaluation for Safety and Multilingual UnderstandingHaneul Yoo, Yongjin Yang, Hwaran LeeACL 2025 · 27 citations
- CrowS-Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language ModelsNikita Nangia, Clara Vania, Rasika Bhalerao, Samuel R. BowmanEMNLP 2020 · 19 citations
Related papers
- KazMMLU: Evaluating Language Models on Kazakh, Russian, and Regional Knowledge of KazakhstanMukhammed Togmanov, Nurdaulet Mukhituly, Diana Turmakhan, Jonibek Mansurov et al.ACL 2025 · 10 citations
- StereoSet: Measuring stereotypical bias in pretrained language modelsMoin Nadeem, Anna Bethke, Siva ReddyACL 2021
- PakBBQ: A Culturally Adapted Bias Benchmark for QAAbdullah Hashmat, Muhammad Arham Mirza, Agha Ali RazaEMNLP 2025
- Who is better at math, Jenny or Jingzhen? Uncovering Stereotypes in Large Language ModelsZara Siddique, Liam D. Turner, Luis Espinosa AnkeEMNLP 2024 · 2 citations
- Instruction Tuning on Public Government and Cultural Data for Low-Resource Language: a Case Study in KazakhNurkhan Laiyk, Daniil Orel, Rituraj Joshi, Maiya Goloburda et al.ACL 2025
