A Survey on Zero Pronoun Translation
Longyue Wang, Siyou Liu, Mingzhou Xu, Linfeng Song, Shuming Shi, Zhaopeng Tu
摘要
Zero pronouns (ZPs) are frequently omitted in pro-drop languages (e.g. Chinese, Hungarian, and Hindi), but should be recalled in non-pro-drop languages (e.g. English). This phenomenon has been studied extensively in machine translation (MT), as it poses a significant challenge for MT systems due to the difficulty in determining the correct antecedent for the pronoun. This survey paper highlights the major works that have been undertaken in zero pronoun translation (ZPT) after the neural revolution so that researchers can recognize the current state and future directions of this field. We provide an organization of the literature based on evolution, dataset, method, and evaluation. In addition, we compare and analyze competing models and evaluation metrics on different benchmarks. We uncover a number of insightful findings such as: 1) ZPT is in line with the development trend of large language model; 2) data limitation causes learning bias in languages and domains; 3) performance improvements are often reported on single benchmarks, but advanced methods are still far from real-world use; 4) general-purpose metrics are not reliable on nuances and complexities of ZPT, emphasizing the necessity of targeted metrics; 5) apart from commonly-cited errors, ZPs will cause risks of gender bias.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- MT2: Towards a Multi-Task Machine Translation Model with Translation-Specific In-Context LearningChunyou Li, Mingtong Liu, Hongxiao Zhang, Yufeng Chen 等EMNLP 2023 · 被引用 3 次
- What the Harm? Quantifying the Tangible Impact of Gender Bias in Machine Translation with a Human-centered StudyBeatrice Savoldi, Sara Papi, Matteo Negri, Ana Guerberof Arenas 等EMNLP 2024 · 被引用 1 次
它引用的顶会 Paper5
- Document-Level Machine Translation with Large Language ModelsLongyue Wang, Chenyang Lyu, Tianbo Ji, Zhirui Zhang 等EMNLP 2023 · 被引用 129 次
- Coupling Context Modeling with Zero Pronoun Recovering for Document-Level Natural Language GenerationXin Tan, Longyin Zhang, Guodong ZhouEMNLP 2021 · 被引用 6 次
- COMET: A Neural Framework for MT EvaluationRicardo Rei, Craig Stewart, Ana C. Farinha, Alon LavieEMNLP 2020 · 被引用 6 次
- GuoFeng: A Benchmark for Zero Pronoun Recovery and TranslationMingzhou Xu, Longyue Wang, Derek F. Wong, Hongye Liu 等EMNLP 2022 · 被引用 5 次
- Pronoun-Targeted Fine-tuning for NMT with Hybrid LossesPrathyusha Jwalapuram, Shafiq R. Joty, Youlin ShenEMNLP 2020 · 被引用 4 次
相关 Paper
- PROSE: A Pronoun Omission Solution for Chinese-English Spoken Language TranslationKe Wang, Xiutian Zhao, Yanghui Li, Wei PengEMNLP 2023 · 被引用 2 次
- What about "em"? How Commercial Machine Translation Fails to Handle (Neo-)PronounsAnne Lauscher, Debora Nozza, Ehm Miltersen, Archie Crowley 等ACL 2023 · 被引用 7 次
- Type B Reflexivization as an Unambiguous Testbed for Multilingual Multi-Task Gender BiasAna Valeria González-Garduño, Maria Barrett, Rasmus Hvingelby, Kellie Webster 等EMNLP 2020
- MORPHOGEN: A Multilingual Benchmark for Evaluating Gender-Aware Morphological GenerationMehul Agarwal, Aditya Aggarwal, Arnav Goel, Medha Hira 等ACL 2026
- HalOmi: A Manually Annotated Benchmark for Multilingual Hallucination and Omission Detection in Machine TranslationDavid Dale, Elena Voita, Janice Lam, Prangthip Hansanti 等EMNLP 2023 · 被引用 8 次
