Investigating the Capabilities and Limitations of Machine Learning for Identifying Bias in English Language Data with Information and Heritage Professionals
Lucy Havens, Benjamin Bach, Melissa Terras, Beatrice Alex
Abstract
Despite numerous efforts to mitigate their biases, ML systems continue to harm already-marginalized people. While predominant ML approaches assume bias can be removed and fair models can be created, we show that these are not always possible, nor desirable, goals. We reframe the problem of ML bias by creating models to identify biased language, drawing attention to a dataset's biases rather than trying to remove them. Then, through a workshop, we evaluated the models for a specific use case: workflows of information and heritage professionals. Our findings demonstrate the limitations of ML for identifying bias due to its contextual nature, the way in which approaches to mitigating it can simultaneously privilege and oppress different communities, and its inevitability. We demonstrate the need to expand ML approaches to bias and fairness, providing a mixed-methods approach to investigating the feasibility of removing bias or achieving fairness in a given ML use case.
• Human-centered computing → HCI design and evaluation methods; • Social and professional topics → Socio-technical systems; • Applied computing → Digital libraries and archives; • Computing methodologies → Machine learning; Natural language processing.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e73bee7c-35c6-4f0d-9a61-33c0eb7fc078Builds on10
- Questioning the AI: Informing Design Practices for Explainable AI User ExperiencesQ. Vera Liao, Daniel M. Gruen, Sarah MillerCHI 2020 · 758 citations
- How WEIRD is CHI?Sebastian Linxen, Christian Sturm, Florian Brühlmann, Vincent Cassau et al.CHI 2021 · 254 citations
- Toward a Perspectivist Turn in Ground Truthing for Predictive ComputingFederico Cabitza, Andrea Campagner, Valerio BasileAAAI 2023 · 236 citations
- Language (Technology) is Power: A Critical Survey of "Bias" in NLPSu Lin Blodgett, Solon Barocas, Hal Daumé III, Hanna M. WallachACL 2020 · 68 citations
- Upstream Mitigation Is Not All You Need: Testing the Bias Transfer Hypothesis in Pre-Trained Language ModelsRyan Steed, Swetasudha Panda, Ari Kobren, Michael L. WickACL 2022 · 52 citations
Related papers
- Don't Erase, Inform! Detecting and Contextualizing Harmful Language in Cultural Heritage CollectionsOrfeas Menis-Mastromichalakis, Jason Liartis, Kristina Rose, Antoine Isaac et al.ACL 2025 · 1 citation
- Fairway: a way to build fair ML softwareJoymallya Chakraborty, Suvodeep Majumder, Zhe Yu, Tim MenziesFSE 2020 · 131 citations
- Social Bias in Multilingual Language Models: A SurveyLance Calvin Lim Gamboa, Yue Feng, Mark G. LeeEMNLP 2025
- A Multilingual, Culture-First Approach to Addressing Misgendering in LLM ApplicationsSunayana Sitaram, Adrian de Wynter, Isobel McCrum, Qilong Gu et al.EMNLP 2025
- Understanding User Sensemaking in Machine Learning Fairness Assessment SystemsZiwei Gu, Jing Nathan Yan, Jeffrey M. RzeszotarskiWWW 2021 · 12 citations
