FairQE: Multi-Agent Framework for Mitigating Gender Bias in Translation Quality Estimation
Jinhee Jang, Juhwan Choi, Dongjin Lee, Seunguk Yu, YoungBin Kim
摘要
Quality Estimation (QE) aims to assess machine translation quality without reference translations, but recent studies have shown that existing QE models exhibit systematic gender bias. In particular, they tend to favor masculine realizations in gender-ambiguous contexts and may assign higher scores to gendermisaligned translations even when gender is explicitly specified. To address these issues, we propose FairQE, a multi-agent-based, fairnessaware QE framework that mitigates gender bias in both gender-ambiguous and genderexplicit scenarios. FairQE detects gender cues, generates gender-flipped translation variants, and combines conventional QE scores with LLM-based bias-mitigating reasoning through a dynamic bias-aware aggregation mechanism. This design preserves the strengths of existing QE models while calibrating their genderrelated biases in a plug-and-play manner. Extensive experiments across multiple gender bias evaluation settings demonstrate that FairQE consistently improves gender fairness over strong QE baselines. Moreover, under MQMbased meta-evaluation following the WMT 2023 Metrics Shared Task, FairQE achieves competitive or improved general QE performance. These results show that gender bias in QE can be effectively mitigated without sacrificing evaluation accuracy, enabling fairer and more reliable translation evaluation. * https://platform.openai.com/docs/ models/gpt-4.1-mini * accuracy (addition, mistranslation, omission, untranslated text) * fluency (character encoding, grammar, inconsistency, punctuation, register, spelling) * locale convention (currency, date, name, telephone, time format) * style (awkward) * terminology (inappropriate for context, inconsistent use) * non-translation * other * or no-error -For EACH identified error, assign a severity level: * Critical * Major * Minor Scoring: -Start from a score of 100 points. -Deduct points as follows: * Critical: -15 points * Major: -5 points * Minor: -1 point -The final score must be between 0 and 100.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- Humans or LLMs as the Judge? A Study on Judgement BiasGuiming Hardy Chen, Shunian Chen, Ziche Liu, Feng Jiang 等EMNLP 2024 · 被引用 37 次
- MT-GenEval: A Counterfactual and Contextual Dataset for Evaluating Gender Accuracy in Machine TranslationAnna Currey, Maria Nadejde, Raghavendra Reddy Pappagari, Mia Mayer 等EMNLP 2022 · 被引用 22 次
- Physician Detection of Clinical Harm in Machine Translation: Quality Estimation Aids in Reliance and Backtranslation Identifies Critical ErrorsNikita Mehandru, Sweta Agrawal, Yimin Xiao, Ge Gao 等EMNLP 2023 · 被引用 9 次
- Watching the Watchers: Exposing Gender Disparities in Machine Translation Quality EstimationEmmanouil Zaranis, Giuseppe Attanasio, Sweta Agrawal, André F. T. MartinsACL 2025 · 被引用 8 次
相关 Paper
- Bias Mitigation in Machine Translation Quality EstimationHanna Behnke, Marina Fomicheva, Lucia SpeciaACL 2022
- Self-Supervised Quality Estimation for Machine TranslationYuanhang Zheng, Zhixing Tan, Meng Zhang, Mieradilijiang Maimaiti 等EMNLP 2021 · 被引用 5 次
- Target-Agnostic Gender-Aware Contrastive Learning for Mitigating Bias in Multilingual Machine TranslationMinwoo Lee, Hyukhun Koh, Kang-il Lee, Dongdong Zhang 等EMNLP 2023 · 被引用 2 次
- DirectQE: Direct Pretraining for Machine Translation Quality EstimationQu Cui, Shujian Huang, Jiahuan Li, Xiang Geng 等AAAI 2021 · 被引用 24 次
- Gender Biases in Automatic Evaluation Metrics for Image CaptioningHaoyi Qiu, Zi-Yi Dou, Tianlu Wang, Asli Celikyilmaz 等EMNLP 2023 · 被引用 6 次
