Scalable and Culturally Specific Stereotype Dataset Construction via Human-LLM Collaboration
Weicheng Ma, John J. Guerrerio, Soroush Vosoughi
Abstract
Warning: This paper contains examples of potentially offensive content. Research on stereotypes in large language models (LLMs) has largely focused on Englishspeaking contexts, due to the lack of datasets in other languages and the high cost of manual annotation in underrepresented cultures. To address this gap, we introduce a cost-efficient human-LLM collaborative annotation framework and apply it to construct EspanStereo, a Spanish-language stereotype dataset spanning multiple Spanish-speaking countries across Europe and Latin America. EspanStereo captures both well-documented stereotypes from prior literature and culturally specific biases absent from English-centric resources. Using LLMs to generate candidate stereotypes and in-culture annotators to validate them, we demonstrate the framework's effectiveness in identifying nuanced, region-specific biases. Our evaluation of Spanish-supporting LLMs using EspanStereo reveals significant variation in stereotypical behavior across countries, highlighting the need for more culturally grounded assessments. Beyond Spanish, our framework is adaptable to other languages and regions, offering a scalable path toward multilingual stereotype benchmarks. This work broadens the scope of stereotype analysis in LLMs and lays the groundwork for comprehensive cross-cultural bias evaluation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 292bfe68-0377-4b43-b1ba-24e7f894c8aeCited by top-tier papers1
Ask how each one uses itBuilds on7
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- French CrowS-Pairs: Extending a challenge dataset for measuring social bias in masked language models to a language other than EnglishAurélie Névéol, Yoann Dupont, Julien Bezançon, Karën FortACL 2022 · 61 citations
- WinoQueer: A Community-in-the-Loop Benchmark for Anti-LGBTQ+ Bias in Large Language ModelsVirginia K. Felkner, Ho-Chun Herbert Chang, Eugene Jang, Jonathan MayACL 2023 · 46 citations
- CrowS-Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language ModelsNikita Nangia, Clara Vania, Rasika Bhalerao, Samuel R. BowmanEMNLP 2020 · 19 citations
- SeeGULL: A Stereotype Benchmark with Broad Geo-Cultural Coverage Leveraging Generative ModelsAkshita Jha, Aida Mostafazadeh Davani, Chandan K. Reddy, Shachi Dave et al.ACL 2023 · 13 citations
Related papers
- Who is better at math, Jenny or Jingzhen? Uncovering Stereotypes in Large Language ModelsZara Siddique, Liam D. Turner, Luis Espinosa AnkeEMNLP 2024 · 2 citations
- EuroGEST: Investigating gender stereotypes in multilingual language modelsJacqueline Rowe, Mateusz Klimaszewski, Liane Guillou, Shannon Vallor et al.EMNLP 2025
- Large Language Models Discriminate Against Speakers of German DialectsMinh Duc Bui, Carolin Holtermann, Valentin Hofmann, Anne Lauscher et al.EMNLP 2025
- StereoSet: Measuring stereotypical bias in pretrained language modelsMoin Nadeem, Anna Bethke, Siva ReddyACL 2021
- HESEIA: A community-based dataset for evaluating social biases in large language models, co-designed in real school settings in Latin AmericaGuido Ivetta, Marcos J. Gomez, Sofía Martinelli, Pietro Palombini et al.EMNLP 2025
