MORPHOGEN: A Multilingual Benchmark for Evaluating Gender-Aware Morphological Generation
Mehul Agarwal, Aditya Aggarwal, Arnav Goel, Medha Hira, Anubha Gupta
摘要
While multilingual large language models (LLMs) perform well on high-level tasks like translation and question answering, their ability to handle grammatical gender and morphological agreement remains underexplored. In morphologically rich languages, gender influences verb conjugation, pronouns, and even first-person constructions with explicit and implicit mentions to gender. We thus introduce MORPHOGEN a morphologically grounded largescale benchmark dataset for evaluating genderaware generation in three typologically diverse grammatically gendered languages i.e. French, Arabic and Hindi. The core task, GENFORM, requires models to rewrite a first-person sentence in the opposite gender while preserving its meaning and structure. We construct a highquality synthetic dataset spanning French, Arabic, and Hindi, and benchmark 15 popular multilingual LLMs (2B-70B) on their ability to perform this transformation. Our results reveal gaps and interesting insights into the handling of morphological gender in current models. MORPHOGEN offers a focused diagnostic lens for gender-aware language modeling and lays the groundwork for future research on inclusive and morphology-sensitive NLP.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- XTREME: A Massively Multilingual Multi-task Benchmark for Evaluating Cross-lingual GeneralisationJunjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig 等ICML 2020 · 被引用 1,132 次
- "I'm sorry to hear that": Finding New Biases in Language Models with a Holistic Descriptor DatasetEric Michael Smith, Melissa Hall, Melanie Kambadur, Eleonora Presani 等EMNLP 2022 · 被引用 56 次
- Gender in Danger? Evaluating Speech Translation Technology on the MuST-SHE CorpusLuisa Bentivogli, Beatrice Savoldi, Matteo Negri, Mattia Antonino Di Gangi 等ACL 2020 · 被引用 40 次
- MT-GenEval: A Counterfactual and Contextual Dataset for Evaluating Gender Accuracy in Machine TranslationAnna Currey, Maria Nadejde, Raghavendra Reddy Pappagari, Mia Mayer 等EMNLP 2022 · 被引用 22 次
- GenderCARE: A Comprehensive Framework for Assessing and Reducing Gender Bias in Large Language ModelsKunsheng Tang, Wenbo Zhou, Jie Zhang, Aishan Liu 等CCS 2024 · 被引用 7 次
相关 Paper
- EuroGEST: Investigating gender stereotypes in multilingual language modelsJacqueline Rowe, Mateusz Klimaszewski, Liane Guillou, Shannon Vallor 等EMNLP 2025
- Mind the Inclusivity Gap: Multilingual Gender-Neutral Translation Evaluation with mGeNTEBeatrice Savoldi, Giuseppe Attanasio, Eleonora Cupin, Eleni Gkovedarou 等EMNLP 2025
- Multilingual Text-to-Image Generation Magnifies Gender StereotypesFelix Friedrich, Katharina Hämmerl, Patrick Schramowski, Manuel Brack 等ACL 2025
- IndicGenBench: A Multilingual Benchmark to Evaluate Generation Capabilities of LLMs on Indic LanguagesHarman Singh, Nitish Gupta, Shikhar Bharadwaj, Dinesh Tewari 等ACL 2024
- GenderAlign: An Alignment Dataset for Mitigating Gender Bias in Large Language ModelsTao Zhang, Ziqian Zeng, Yuxiang Xiao, Huiping Zhuang 等ACL 2025 · 被引用 18 次
