Don't Erase, Inform! Detecting and Contextualizing Harmful Language in Cultural Heritage Collections
Orfeas Menis-Mastromichalakis, Jason Liartis, Kristina Rose, Antoine Isaac, Giorgos Stamou
Abstract
Cultural Heritage (CH) data hold invaluable knowledge, reflecting the history, traditions, and identities of societies, and shaping our understanding of the past and present. However, many CH collections contain outdated or offensive descriptions that reflect historical biases. CH Institutions (CHIs) face significant challenges in curating these data due to the vast scale and complexity of the task. To address this, we develop an AI-powered tool that detects offensive terms in CH metadata and provides contextual insights into their historical background and contemporary perception. We leverage a multilingual vocabulary co-created with marginalized communities, researchers, and CH professionals, along with traditional NLP techniques and Large Language Models (LLMs). Available as a standalone web app and integrated with major CH platforms, the tool has processed over 7.9 million records, contextualizing the contentious terms detected in their metadata. Rather than erasing these terms, our approach seeks to inform, making biases visible and providing actionable insights for creating more inclusive and accessible CH collections.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on2
- Learning The Difference That Makes A Difference With Counterfactually-Augmented DataDivyansh Kaushik, Eduard H. Hovy, Zachary Chase LiptonICLR 2020 · 625 citations
- Demographics Should Not Be the Reason of Toxicity: Mitigating Discrimination in Text Classifications with Instance WeightingGuanhua Zhang, Bing Bai, Junqi Zhang, Kun Bai et al.ACL 2020 · 56 citations
Related papers
- Investigating the Capabilities and Limitations of Machine Learning for Identifying Bias in English Language Data with Information and Heritage ProfessionalsLucy Havens, Benjamin Bach, Melissa Terras, Beatrice AlexCHI 2025 · 3 citations
- Identifying, Explaining, and Correcting Ableist Language with AIKynnedy Simone Smith, Lydia B. Chilton, Danielle BraggCHI 2026 · 1 citation
- Integrating Machine Learning Data with Symbolic Knowledge from Collaboration Practices of Curators to Improve Conversational SystemsClaudio Santos Pinhanez, Heloisa Candello, Paulo Rodrigo Cavalin, Mauro Carlos Pichiliani et al.CHI 2021 · 7 citations
- Generative AI and Perceptual Harms: Who's Suspected of using LLMs?Kowe Kadoma, Danaë Metaxa, Mor NaamanCHI 2025 · 28 citations
- "I'm sorry to hear that": Finding New Biases in Language Models with a Holistic Descriptor DatasetEric Michael Smith, Melissa Hall, Melanie Kambadur, Eleonora Presani et al.EMNLP 2022 · 56 citations
