Investigating Failures of Automatic Translationin the Case of Unambiguous Gender
Adi Renduchintala, Adina Williams
Abstract
Transformer based models are the modern work horses for neural machine translation (NMT), reaching state of the art across several benchmarks. Despite their impressive accuracy, we observe a systemic and rudimentary class of errors made by transformer based models with regards to translating from a language that doesn't mark gender on nouns into others that do. We find that even when the surrounding context provides unambiguous evidence of the appropriate grammatical gender marking, no transformer based model we tested was able to accurately gender occupation nouns systematically. We release an evaluation scheme and dataset for measuring the ability of transformer based NMT models to translate gender morphology correctly in unambiguous contexts across syntactically diverse sentences. Our dataset translates from an English source into 20 languages from several different language families. With the availability of this dataset, our hope is that the NMT community can iterate on solutions for this class of especially egregious errors. Source/Target Label Src: My sister is a carpenter 4 . Correct Tgt: Mi hermana es carpenteria(f) 4 . Src: That nurse 1 is a funny man . Wrong Tgt: Esa enfermera(f) 1 es un tipo gracioso . Src: The engineer 1 is her emotional mother . Inconclusive Tgt: La ingeniería(?) 1 es su madre emocional .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5eff54af-598e-4d70-ab43-1a0441e71057Cited by top-tier papers8
- Dynaboard: An Evaluation-As-A-Service Platform for Holistic Next-Generation BenchmarkingZhiyi Ma, Kawin Ethayarajh, Tristan Thrush, Somya Jain et al.NeurIPS 2021 · 76 citations
- Perturbation Augmentation for Fairer NLPRebecca Qian, Candace Ross, Jude Fernandes, Eric Michael Smith et al.EMNLP 2022 · 54 citations
- Contrastive Conditioning for Assessing Disambiguation in MT: A Case Study of Distilled BiasJannis Vamvas, Rico SennrichEMNLP 2021 · 12 citations
- Exploiting Biased Models to De-bias Text: A Gender-Fair Rewriting ModelChantal Amrhein, Florian Schottmann, Rico Sennrich, Samuel LäubliACL 2023 · 7 citations
- Hi Guys or Hi Folks? Benchmarking Gender-Neutral Machine Translation with the GeNTE CorpusAndrea Piergentili, Beatrice Savoldi, Dennis Fucci, Matteo Negri et al.EMNLP 2023 · 3 citations
Builds on3
- Queens are Powerful too: Mitigating Gender Bias in Dialogue GenerationEmily Dinan, Angela Fan, Adina Williams, Jack Urbanek et al.EMNLP 2020 · 14 citations
- Multi-Dimensional Gender Bias ClassificationEmily Dinan, Angela Fan, Ledell Wu, Jason Weston et al.EMNLP 2020 · 7 citations
- Type B Reflexivization as an Unambiguous Testbed for Multilingual Multi-Task Gender BiasAna Valeria González-Garduño, Maria Barrett, Rasmus Hvingelby, Kellie Webster et al.EMNLP 2020
Related papers
- Measuring and Mitigating Name Biases in Neural Machine TranslationJun Wang, Benjamin I. P. Rubinstein, Trevor CohnACL 2022 · 31 citations
- MT-GenEval: A Counterfactual and Contextual Dataset for Evaluating Gender Accuracy in Machine TranslationAnna Currey, Maria Nadejde, Raghavendra Reddy Pappagari, Mia Mayer et al.EMNLP 2022 · 22 citations
- Assumed Identities: Quantifying Gender Bias in Machine Translation of Gender-Ambiguous Occupational TermsOrfeas Menis-Mastromichalakis, Giorgos Filandrianos, Maria Symeonaki, Giorgos StamouEMNLP 2025
- Mind the Inclusivity Gap: Multilingual Gender-Neutral Translation Evaluation with mGeNTEBeatrice Savoldi, Giuseppe Attanasio, Eleonora Cupin, Eleni Gkovedarou et al.EMNLP 2025
- GFST: Gender-Filtered Self-Training for More Accurate Gender in TranslationPrafulla Kumar Choubey, Anna Currey, Prashant Mathur, Georgiana DinuEMNLP 2021 · 7 citations
