GFST: Gender-Filtered Self-Training for More Accurate Gender in Translation
Prafulla Kumar Choubey, Anna Currey, Prashant Mathur, Georgiana Dinu
Abstract
Targeted evaluations have found that machine translation systems often output incorrect gender in translations, even when the gender is clear from context. Furthermore, these incorrectly gendered translations have the potential to reflect or amplify social biases. We propose gender-filtered self-training (GFST) to improve gender translation accuracy on unambiguously gendered inputs. Our GFST approach uses a source monolingual corpus and an initial model to generate gender-specific pseudo-parallel corpora which are then filtered and added to the training data. We evaluate GFST on translation from English into five languages, finding that it improves gender accuracy without damaging generic quality. We also show the viability of GFST on several experimental settings, including re-training from scratch, fine-tuning, controlling the gender balance of the data, forward translation, and back-translation. 1 * Equal contribution. † Work done as an intern at Amazon AI Translate.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a223a6f2-0a1a-44bd-9bb9-7b936efe0d95Cited by top-tier papers3
- MT-GenEval: A Counterfactual and Contextual Dataset for Evaluating Gender Accuracy in Machine TranslationAnna Currey, Maria Nadejde, Raghavendra Reddy Pappagari, Mia Mayer et al.EMNLP 2022 · 22 citations
- MISGENDERED: Limits of Large Language Models in Understanding PronounsTamanna Hossain, Sunipa Dev, Sameer SinghACL 2023 · 8 citations
- Target-Agnostic Gender-Aware Contrastive Learning for Mitigating Bias in Multilingual Machine TranslationMinwoo Lee, Hyukhun Koh, Kang-il Lee, Dongdong Zhang et al.EMNLP 2023 · 2 citations
Builds on5
- Predictive Biases in Natural Language Processing Models: A Conceptual Framework and OverviewDeven Shah, H. Andrew Schwartz, Dirk HovyACL 2020 · 93 citations
- Language (Technology) is Power: A Critical Survey of "Bias" in NLPSu Lin Blodgett, Solon Barocas, Hal Daumé III, Hanna M. WallachACL 2020 · 68 citations
- Gender in Danger? Evaluating Speech Translation Technology on the MuST-SHE CorpusLuisa Bentivogli, Beatrice Savoldi, Matteo Negri, Mattia Antonino Di Gangi et al.ACL 2020 · 40 citations
- Distilling Multiple Domains for Neural Machine TranslationAnna Currey, Prashant Mathur, Georgiana DinuEMNLP 2020 · 19 citations
- Reducing Gender Bias in Neural Machine Translation as a Domain Adaptation ProblemDanielle Saunders, Bill ByrneACL 2020 · 7 citations
Related papers
- Exploiting Biased Models to De-bias Text: A Gender-Fair Rewriting ModelChantal Amrhein, Florian Schottmann, Rico Sennrich, Samuel LäubliACL 2023 · 7 citations
- Measuring and Mitigating Name Biases in Neural Machine TranslationJun Wang, Benjamin I. P. Rubinstein, Trevor CohnACL 2022 · 31 citations
- A Tale of Pronouns: Interpretability Informs Gender Bias Mitigation for Fairer Instruction-Tuned Machine TranslationGiuseppe Attanasio, Flor Miriam Plaza del Arco, Debora Nozza, Anne LauscherEMNLP 2023 · 3 citations
- Investigating Failures of Automatic Translationin the Case of Unambiguous GenderAdi Renduchintala, Adina WilliamsACL 2022 · 28 citations
- What the Harm? Quantifying the Tangible Impact of Gender Bias in Machine Translation with a Human-centered StudyBeatrice Savoldi, Sara Papi, Matteo Negri, Ana Guerberof Arenas et al.EMNLP 2024 · 1 citation
