FairQE: Multi-Agent Framework for Mitigating Gender Bias in Translation Quality Estimation
Jinhee Jang, Juhwan Choi, Dongjin Lee, Seunguk Yu, YoungBin Kim
Abstract
Quality Estimation (QE) aims to assess machine translation quality without reference translations, but recent studies have shown that existing QE models exhibit systematic gender bias. In particular, they tend to favor masculine realizations in gender-ambiguous contexts and may assign higher scores to gendermisaligned translations even when gender is explicitly specified. To address these issues, we propose FairQE, a multi-agent-based, fairnessaware QE framework that mitigates gender bias in both gender-ambiguous and genderexplicit scenarios. FairQE detects gender cues, generates gender-flipped translation variants, and combines conventional QE scores with LLM-based bias-mitigating reasoning through a dynamic bias-aware aggregation mechanism. This design preserves the strengths of existing QE models while calibrating their genderrelated biases in a plug-and-play manner. Extensive experiments across multiple gender bias evaluation settings demonstrate that FairQE consistently improves gender fairness over strong QE baselines. Moreover, under MQMbased meta-evaluation following the WMT 2023 Metrics Shared Task, FairQE achieves competitive or improved general QE performance. These results show that gender bias in QE can be effectively mitigated without sacrificing evaluation accuracy, enabling fairer and more reliable translation evaluation. * https://platform.openai.com/docs/ models/gpt-4.1-mini * accuracy (addition, mistranslation, omission, untranslated text) * fluency (character encoding, grammar, inconsistency, punctuation, register, spelling) * locale convention (currency, date, name, telephone, time format) * style (awkward) * terminology (inappropriate for context, inconsistent use) * non-translation * other * or no-error -For EACH identified error, assign a severity level: * Critical * Major * Minor Scoring: -Start from a score of 100 points. -Deduct points as follows: * Critical: -15 points * Major: -5 points * Minor: -1 point -The final score must be between 0 and 100.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 110e52f3-d332-4c7d-946f-e39af99f290aBuilds on12
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Humans or LLMs as the Judge? A Study on Judgement BiasGuiming Hardy Chen, Shunian Chen, Ziche Liu, Feng Jiang et al.EMNLP 2024 · 37 citations
- MT-GenEval: A Counterfactual and Contextual Dataset for Evaluating Gender Accuracy in Machine TranslationAnna Currey, Maria Nadejde, Raghavendra Reddy Pappagari, Mia Mayer et al.EMNLP 2022 · 22 citations
- Physician Detection of Clinical Harm in Machine Translation: Quality Estimation Aids in Reliance and Backtranslation Identifies Critical ErrorsNikita Mehandru, Sweta Agrawal, Yimin Xiao, Ge Gao et al.EMNLP 2023 · 9 citations
- Watching the Watchers: Exposing Gender Disparities in Machine Translation Quality EstimationEmmanouil Zaranis, Giuseppe Attanasio, Sweta Agrawal, André F. T. MartinsACL 2025 · 8 citations
Related papers
- Bias Mitigation in Machine Translation Quality EstimationHanna Behnke, Marina Fomicheva, Lucia SpeciaACL 2022
- Self-Supervised Quality Estimation for Machine TranslationYuanhang Zheng, Zhixing Tan, Meng Zhang, Mieradilijiang Maimaiti et al.EMNLP 2021 · 5 citations
- Target-Agnostic Gender-Aware Contrastive Learning for Mitigating Bias in Multilingual Machine TranslationMinwoo Lee, Hyukhun Koh, Kang-il Lee, Dongdong Zhang et al.EMNLP 2023 · 2 citations
- DirectQE: Direct Pretraining for Machine Translation Quality EstimationQu Cui, Shujian Huang, Jiahuan Li, Xiang Geng et al.AAAI 2021 · 24 citations
- Gender Biases in Automatic Evaluation Metrics for Image CaptioningHaoyi Qiu, Zi-Yi Dou, Tianlu Wang, Asli Celikyilmaz et al.EMNLP 2023 · 6 citations
