Controlled Analyses of Social Biases in Wikipedia Bios
Anjalie Field, Chan Young Park, Kevin Z. Lin, Yulia Tsvetkov
Abstract
Social biases on Wikipedia, a widely-read global platform, could greatly influence public opinion. While prior research has examined man/woman gender bias in biography articles, possible influences of other demographic attributes limit conclusions. In this work, we present a methodology for analyzing Wikipedia pages about people that isolates dimensions of interest (e.g., gender), from other attributes (e.g., occupation). Given a target corpus for analysis (e.g. biographies about women), we present a method for constructing a comparison corpus that matches the target corpus in as many attributes as possible, except the target one. We develop evaluation metrics to measure how well the comparison corpus aligns with the target corpus and then examine how articles about gender and racial minorities (cis. women, non-binary people, transgender women, and transgender men; African American, Asian American, and Hispanic/Latinx American people) differ from other articles. In addition to identifying suspect social biases, our results show that failing to control for covariates can result in different conclusions and veil biases. Our contributions include methodology that facilitates further analyses of bias in Wikipedia articles, findings that can aid Wikipedia editors in reducing biases, and a framework and evaluation metrics to guide future work in this area. CCS CONCEPTS • Human-centered computing → Empirical studies in collaborative and social computing; • Computing methodologies → Natural language processing; • Information systems → Wikis.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Gendered Mental Health Stigma in Masked Language ModelsInna W. Lin, Lucille Njoo, Anjalie Field, Ashish Sharma et al.EMNLP 2022 · 15 citations
- White Men Lead, Black Women Help? Benchmarking and Mitigating Language Agency Social Biases in LLMsYixin Wan, Kai-Wei ChangACL 2025 · 9 citations
- Locating Information Gaps and Narrative Inconsistencies Across Languages: A Case Study of LGBT People Portrayals on WikipediaFarhan Samir, Chan Young Park, Anjalie Field, Vered Shwartz et al.EMNLP 2024 · 2 citations
Builds on3
- Language (Technology) is Power: A Critical Survey of "Bias" in NLPSu Lin Blodgett, Solon Barocas, Hal Daumé III, Hanna M. WallachACL 2020 · 68 citations
- Text and Causal Inference: A Review of Using Text to Remove Confounding from Causal EstimatesKatherine A. Keith, David D. Jensen, Brendan O'ConnorACL 2020 · 16 citations
- A Survey of Race, Racism, and Anti-Racism in NLPAnjalie Field, Su Lin Blodgett, Zeerak Waseem, Yulia TsvetkovACL 2021
Related papers
- WikiBio: a Semantic Resource for the Intersectional Analysis of Biographical EventsMarco Antonio Stranisci, Rossana Damiano, Enrico Mensa, Viviana Patti et al.ACL 2023 · 2 citations
- A General Framework for Implicit and Explicit Debiasing of Distributional Word Vector SpacesAnne Lauscher, Goran Glavas, Simone Paolo Ponzetto, Ivan VulicAAAI 2020 · 68 citations
- Toward Gender-Inclusive Coreference ResolutionYang Trista Cao, Hal Daumé IIIACL 2020 · 20 citations
- Social Bias in Multilingual Language Models: A SurveyLance Calvin Lim Gamboa, Yue Feng, Mark G. LeeEMNLP 2025
- Mitigating Language-Dependent Ethnic Bias in BERTJaimeen Ahn, Alice OhEMNLP 2021 · 5 citations
