Watching the Watchers: Exposing Gender Disparities in Machine Translation Quality Estimation
Emmanouil Zaranis, Giuseppe Attanasio, Sweta Agrawal, André F. T. Martins
Abstract
Quality estimation (QE)-the automatic assessment of translation quality-has recently become crucial across several stages of the translation pipeline, from data curation to training and decoding. While QE metrics have been optimized to align with human judgments, whether they encode social biases has been largely overlooked. Biased QE risks favoring certain demographic groups over others, e.g., by exacerbating gaps in visibility and usability. This paper defines and investigates gender bias of QE metrics and discusses its downstream implications for machine translation (MT). Experiments with state-ofthe-art QE metrics across multiple domains, datasets, and languages reveal significant bias. When a human entity's gender in the source is undisclosed, masculine-inflected translations score higher than feminine-inflected ones, and gender-neutral translations are penalized. Even when contextual cues disambiguate gender, using context-aware QE metrics leads to more errors in selecting the correct translation inflection for feminine referents than for masculine ones. Moreover, a biased QE metric affects data filtering and quality-aware decoding. Our findings underscore the need for a renewed focus on developing and evaluating QE metrics centered on gender. 1 Tymoshenko has worked as a practicing economist and academic. Tymoshenko ha lavorato come economista e accademica. F Context Source Hypotheses . 83 .89 Tymoshenko ha lavorato come economista e accademico. M .98 .91 Tymoshenko ha lavorato come economista e in università. N .95 QE scores wo/ and w/ context In 1999, she defended her PhD dissertation, titled State Regulation of the tax system, at the Kyiv National Economic University and received a Ph.D. in Economics.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 10a7c6fc-ed95-4381-8591-802d6c4aa496Cited by top-tier papers3
- FairQE: Multi-Agent Framework for Mitigating Gender Bias in Translation Quality EstimationJinhee Jang, Juhwan Choi, Dongjin Lee, Seunguk Yu et al.ACL 2026
- Mind the Inclusivity Gap: Multilingual Gender-Neutral Translation Evaluation with mGeNTEBeatrice Savoldi, Giuseppe Attanasio, Eleonora Cupin, Eleni Gkovedarou et al.EMNLP 2025
- Assumed Identities: Quantifying Gender Bias in Machine Translation of Gender-Ambiguous Occupational TermsOrfeas Menis-Mastromichalakis, Giorgos Filandrianos, Maria Symeonaki, Giorgos StamouEMNLP 2025
Builds on20
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- CLIPScore: A Reference-free Evaluation Metric for Image CaptioningJack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras et al.EMNLP 2021 · 937 citations
- Large Language Models Are Not Robust Multiple Choice SelectorsChujie Zheng, Hao Zhou, Fandong Meng, Jie Zhou et al.ICLR 2024 · 424 citations
- BLEURT: Learning Robust Metrics for Text GenerationThibault Sellam, Dipanjan Das, Ankur P. ParikhACL 2020 · 40 citations
- INSTRUCTSCORE: Towards Explainable Text Generation Evaluation with Automatic FeedbackWenda Xu, Danqing Wang, Liangming Pan, Zhenqiao Song et al.EMNLP 2023 · 36 citations
Related papers
- What the Harm? Quantifying the Tangible Impact of Gender Bias in Machine Translation with a Human-centered StudyBeatrice Savoldi, Sara Papi, Matteo Negri, Ana Guerberof Arenas et al.EMNLP 2024 · 1 citation
- Self-Supervised Quality Estimation for Machine TranslationYuanhang Zheng, Zhixing Tan, Meng Zhang, Mieradilijiang Maimaiti et al.EMNLP 2021 · 5 citations
- Bias Mitigation in Machine Translation Quality EstimationHanna Behnke, Marina Fomicheva, Lucia SpeciaACL 2022
- Improving Translation Quality Estimation with Bias MitigationHui Huang, Shuangzhi Wu, Kehai Chen, Hui Di et al.ACL 2023 · 2 citations
- MT-GenEval: A Counterfactual and Contextual Dataset for Evaluating Gender Accuracy in Machine TranslationAnna Currey, Maria Nadejde, Raghavendra Reddy Pappagari, Mia Mayer et al.EMNLP 2022 · 22 citations
