Investigating the Capabilities and Limitations of Machine Learning for Identifying Bias in English Language Data with Information and Heritage Professionals
Lucy Havens, Benjamin Bach, Melissa Terras, Beatrice Alex
摘要
Despite numerous efforts to mitigate their biases, ML systems continue to harm already-marginalized people. While predominant ML approaches assume bias can be removed and fair models can be created, we show that these are not always possible, nor desirable, goals. We reframe the problem of ML bias by creating models to identify biased language, drawing attention to a dataset's biases rather than trying to remove them. Then, through a workshop, we evaluated the models for a specific use case: workflows of information and heritage professionals. Our findings demonstrate the limitations of ML for identifying bias due to its contextual nature, the way in which approaches to mitigating it can simultaneously privilege and oppress different communities, and its inevitability. We demonstrate the need to expand ML approaches to bias and fairness, providing a mixed-methods approach to investigating the feasibility of removing bias or achieving fairness in a given ML use case.
• Human-centered computing → HCI design and evaluation methods; • Social and professional topics → Socio-technical systems; • Applied computing → Digital libraries and archives; • Computing methodologies → Machine learning; Natural language processing.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- Questioning the AI: Informing Design Practices for Explainable AI User ExperiencesQ. Vera Liao, Daniel M. Gruen, Sarah MillerCHI 2020 · 被引用 758 次
- How WEIRD is CHI?Sebastian Linxen, Christian Sturm, Florian Brühlmann, Vincent Cassau 等CHI 2021 · 被引用 254 次
- Toward a Perspectivist Turn in Ground Truthing for Predictive ComputingFederico Cabitza, Andrea Campagner, Valerio BasileAAAI 2023 · 被引用 236 次
- Language (Technology) is Power: A Critical Survey of "Bias" in NLPSu Lin Blodgett, Solon Barocas, Hal Daumé III, Hanna M. WallachACL 2020 · 被引用 68 次
- Upstream Mitigation Is Not All You Need: Testing the Bias Transfer Hypothesis in Pre-Trained Language ModelsRyan Steed, Swetasudha Panda, Ari Kobren, Michael L. WickACL 2022 · 被引用 52 次
相关 Paper
- Don't Erase, Inform! Detecting and Contextualizing Harmful Language in Cultural Heritage CollectionsOrfeas Menis-Mastromichalakis, Jason Liartis, Kristina Rose, Antoine Isaac 等ACL 2025 · 被引用 1 次
- Fairway: a way to build fair ML softwareJoymallya Chakraborty, Suvodeep Majumder, Zhe Yu, Tim MenziesFSE 2020 · 被引用 131 次
- Social Bias in Multilingual Language Models: A SurveyLance Calvin Lim Gamboa, Yue Feng, Mark G. LeeEMNLP 2025
- A Multilingual, Culture-First Approach to Addressing Misgendering in LLM ApplicationsSunayana Sitaram, Adrian de Wynter, Isobel McCrum, Qilong Gu 等EMNLP 2025
- Understanding User Sensemaking in Machine Learning Fairness Assessment SystemsZiwei Gu, Jing Nathan Yan, Jeffrey M. RzeszotarskiWWW 2021 · 被引用 12 次
