A Tale of Pronouns: Interpretability Informs Gender Bias Mitigation for Fairer Instruction-Tuned Machine Translation
Giuseppe Attanasio, Flor Miriam Plaza del Arco, Debora Nozza, Anne Lauscher
Abstract
Recent instruction fine-tuned models can solve multiple NLP tasks when prompted to do so, with machine translation (MT) being a prominent use case. However, current research often focuses on standard performance benchmarks, leaving compelling fairness and ethical considerations behind. In MT, this might lead to misgendered translations, resulting, among other harms, in the perpetuation of stereotypes and prejudices. In this work, we address this gap by investigating whether and to what extent such models exhibit gender bias in machine translation and how we can mitigate it. Concretely, we compute established gender bias metrics on the WinoMT corpus from English to German and Spanish. We discover that IFT models default to male-inflected translations, even disregarding female occupational stereotypes. Next, using interpretability methods, we unveil that models systematically overlook the pronoun indicating the gender of a target occupation in misgendered translations. Finally, based on this finding, we propose an easy-to-implement and effective bias mitigation solution based on fewshot learning that leads to significantly fairer translations. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4823f811-6cfb-4638-a3a3-9028787c77b7Cited by top-tier papers6
- Watching the Watchers: Exposing Gender Disparities in Machine Translation Quality EstimationEmmanouil Zaranis, Giuseppe Attanasio, Sweta Agrawal, André F. T. MartinsACL 2025 · 8 citations
- Metrics for What, Metrics for Whom: Assessing Actionability of Bias Evaluation Metrics in NLPPieter Delobelle, Giuseppe Attanasio, Debora Nozza, Su Lin Blodgett et al.EMNLP 2024 · 4 citations
- What the Harm? Quantifying the Tangible Impact of Gender Bias in Machine Translation with a Human-centered StudyBeatrice Savoldi, Sara Papi, Matteo Negri, Ana Guerberof Arenas et al.EMNLP 2024 · 1 citation
- The Lou Dataset - Exploring the Impact of Gender-Fair Language in German Text ClassificationAndreas Waldis, Joel Birrer, Anne Lauscher, Iryna GurevychEMNLP 2024 · 1 citation
- Mind the Inclusivity Gap: Multilingual Gender-Neutral Translation Evaluation with mGeNTEBeatrice Savoldi, Giuseppe Attanasio, Eleonora Cupin, Eleni Gkovedarou et al.EMNLP 2025
Builds on15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- Multitask Prompted Training Enables Zero-Shot Task GeneralizationVictor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach et al.ICLR 2022 · 1,976 citations
- The Flan Collection: Designing Data and Methods for Effective Instruction TuningShayne Longpre, Le Hou, Tu Vu, Albert Webson et al.ICML 2023 · 908 citations
Related papers
- Assumed Identities: Quantifying Gender Bias in Machine Translation of Gender-Ambiguous Occupational TermsOrfeas Menis-Mastromichalakis, Giorgos Filandrianos, Maria Symeonaki, Giorgos StamouEMNLP 2025
- Exploiting Biased Models to De-bias Text: A Gender-Fair Rewriting ModelChantal Amrhein, Florian Schottmann, Rico Sennrich, Samuel LäubliACL 2023 · 7 citations
- Measuring and Mitigating Name Biases in Neural Machine TranslationJun Wang, Benjamin I. P. Rubinstein, Trevor CohnACL 2022 · 31 citations
- Reducing Gender Bias in Neural Machine Translation as a Domain Adaptation ProblemDanielle Saunders, Bill ByrneACL 2020 · 7 citations
- GFST: Gender-Filtered Self-Training for More Accurate Gender in TranslationPrafulla Kumar Choubey, Anna Currey, Prashant Mathur, Georgiana DinuEMNLP 2021 · 7 citations
