A Survey on Zero Pronoun Translation
Longyue Wang, Siyou Liu, Mingzhou Xu, Linfeng Song, Shuming Shi, Zhaopeng Tu
Abstract
Zero pronouns (ZPs) are frequently omitted in pro-drop languages (e.g. Chinese, Hungarian, and Hindi), but should be recalled in non-pro-drop languages (e.g. English). This phenomenon has been studied extensively in machine translation (MT), as it poses a significant challenge for MT systems due to the difficulty in determining the correct antecedent for the pronoun. This survey paper highlights the major works that have been undertaken in zero pronoun translation (ZPT) after the neural revolution so that researchers can recognize the current state and future directions of this field. We provide an organization of the literature based on evolution, dataset, method, and evaluation. In addition, we compare and analyze competing models and evaluation metrics on different benchmarks. We uncover a number of insightful findings such as: 1) ZPT is in line with the development trend of large language model; 2) data limitation causes learning bias in languages and domains; 3) performance improvements are often reported on single benchmarks, but advanced methods are still far from real-world use; 4) general-purpose metrics are not reliable on nuances and complexities of ZPT, emphasizing the necessity of targeted metrics; 5) apart from commonly-cited errors, ZPs will cause risks of gender bias.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1293b96f-5b04-4d77-b609-24afbbaffa9aCited by top-tier papers2
- MT2: Towards a Multi-Task Machine Translation Model with Translation-Specific In-Context LearningChunyou Li, Mingtong Liu, Hongxiao Zhang, Yufeng Chen et al.EMNLP 2023 · 3 citations
- What the Harm? Quantifying the Tangible Impact of Gender Bias in Machine Translation with a Human-centered StudyBeatrice Savoldi, Sara Papi, Matteo Negri, Ana Guerberof Arenas et al.EMNLP 2024 · 1 citation
Builds on5
- Document-Level Machine Translation with Large Language ModelsLongyue Wang, Chenyang Lyu, Tianbo Ji, Zhirui Zhang et al.EMNLP 2023 · 129 citations
- Coupling Context Modeling with Zero Pronoun Recovering for Document-Level Natural Language GenerationXin Tan, Longyin Zhang, Guodong ZhouEMNLP 2021 · 6 citations
- COMET: A Neural Framework for MT EvaluationRicardo Rei, Craig Stewart, Ana C. Farinha, Alon LavieEMNLP 2020 · 6 citations
- GuoFeng: A Benchmark for Zero Pronoun Recovery and TranslationMingzhou Xu, Longyue Wang, Derek F. Wong, Hongye Liu et al.EMNLP 2022 · 5 citations
- Pronoun-Targeted Fine-tuning for NMT with Hybrid LossesPrathyusha Jwalapuram, Shafiq R. Joty, Youlin ShenEMNLP 2020 · 4 citations
Related papers
- PROSE: A Pronoun Omission Solution for Chinese-English Spoken Language TranslationKe Wang, Xiutian Zhao, Yanghui Li, Wei PengEMNLP 2023 · 2 citations
- What about "em"? How Commercial Machine Translation Fails to Handle (Neo-)PronounsAnne Lauscher, Debora Nozza, Ehm Miltersen, Archie Crowley et al.ACL 2023 · 7 citations
- Type B Reflexivization as an Unambiguous Testbed for Multilingual Multi-Task Gender BiasAna Valeria González-Garduño, Maria Barrett, Rasmus Hvingelby, Kellie Webster et al.EMNLP 2020
- MORPHOGEN: A Multilingual Benchmark for Evaluating Gender-Aware Morphological GenerationMehul Agarwal, Aditya Aggarwal, Arnav Goel, Medha Hira et al.ACL 2026
- HalOmi: A Manually Annotated Benchmark for Multilingual Hallucination and Omission Detection in Machine TranslationDavid Dale, Elena Voita, Janice Lam, Prangthip Hansanti et al.EMNLP 2023 · 8 citations
