What the Harm? Quantifying the Tangible Impact of Gender Bias in Machine Translation with a Human-centered Study
Beatrice Savoldi, Sara Papi, Matteo Negri, Ana Guerberof Arenas, Luisa Bentivogli
Abstract
Gender bias in machine translation (MT) is recognized as an issue that can harm people and society. And yet, advancements in the field rarely involve people, the final MT users, or inform how they might be impacted by biased technologies. Current evaluations are often restricted to automatic methods, which offer an opaque estimate of what the downstream impact of gender disparities might be. We conduct an extensive human-centered study to examine if and to what extent bias in MT brings harms with tangible costs, such as quality of service gaps across women and men. To this aim, we collect behavioral data from ∼90 participants, who post-edited MT outputs to ensure correct gender translation. Across multiple datasets, languages, and types of users, our study shows that feminine post-editing demands significantly more technical and temporal effort, also corresponding to higher financial costs. Existing bias measurements, however, fail to reflect the found disparities. Our findings advocate for human-centered approaches that can inform the societal impact of bias.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0547d1d2-b8f4-4f75-8457-bbc6c2ec9768Cited by top-tier papers6
- Watching the Watchers: Exposing Gender Disparities in Machine Translation Quality EstimationEmmanouil Zaranis, Giuseppe Attanasio, Sweta Agrawal, André F. T. MartinsACL 2025 · 8 citations
- Exploring the Translation Mechanism of Large Language ModelsHongbin Zhang, Kehai Chen, Xuefeng Bai, Xiucheng Li et al.NeurIPS 2025 · 4 citations
- An Interdisciplinary Approach to Human-Centered Machine TranslationMarine Carpuat, Omri Asscher, Kalika Bali, Luisa Bentivogli et al.EMNLP 2025 · 2 citations
- FairQE: Multi-Agent Framework for Mitigating Gender Bias in Translation Quality EstimationJinhee Jang, Juhwan Choi, Dongjin Lee, Seunguk Yu et al.ACL 2026
- Mind the Inclusivity Gap: Multilingual Gender-Neutral Translation Evaluation with mGeNTEBeatrice Savoldi, Giuseppe Attanasio, Eleonora Cupin, Eleni Gkovedarou et al.EMNLP 2025
Builds on22
- Harms of Gender Exclusivity and Challenges in Non-Binary Representation in Language TechnologiesSunipa Dev, Masoud Monajatipoor, Anaelia Ovalle, Arjun Subramonian et al.EMNLP 2021 · 113 citations
- Measuring and Mitigating Name Biases in Neural Machine TranslationJun Wang, Benjamin I. P. Rubinstein, Trevor CohnACL 2022 · 31 citations
- Under the Morphosyntactic Lens: A Multifaceted Evaluation of Gender Bias in Speech TranslationBeatrice Savoldi, Marco Gaido, Luisa Bentivogli, Matteo Negri et al.ACL 2022 · 30 citations
- Investigating Failures of Automatic Translationin the Case of Unambiguous GenderAdi Renduchintala, Adina WilliamsACL 2022 · 28 citations
- MT-GenEval: A Counterfactual and Contextual Dataset for Evaluating Gender Accuracy in Machine TranslationAnna Currey, Maria Nadejde, Raghavendra Reddy Pappagari, Mia Mayer et al.EMNLP 2022 · 22 citations
Related papers
- A Rose by Any Other Name would not Smell as Sweet: Social Bias in Names MistranslationSandra Sandoval, Jieyu Zhao, Marine Carpuat, Hal Daumé IIIEMNLP 2023 · 2 citations
- GFST: Gender-Filtered Self-Training for More Accurate Gender in TranslationPrafulla Kumar Choubey, Anna Currey, Prashant Mathur, Georgiana DinuEMNLP 2021 · 7 citations
- Physician Detection of Clinical Harm in Machine Translation: Quality Estimation Aids in Reliance and Backtranslation Identifies Critical ErrorsNikita Mehandru, Sweta Agrawal, Yimin Xiao, Ge Gao et al.EMNLP 2023 · 9 citations
- Reducing Gender Bias in Neural Machine Translation as a Domain Adaptation ProblemDanielle Saunders, Bill ByrneACL 2020 · 7 citations
- Investigating the Helpfulness of Word-Level Quality Estimation for Post-Editing Machine Translation OutputRaksha Shenoy, Nico Herbig, Antonio Krüger, Josef van GenabithEMNLP 2021 · 3 citations
