Paraphrasing as Zero-shot Translation with Feature-guided Diversity Enhancement
Ziyue Yan, Hongying Zan, Xinglin Lyu, Hongfei Xu
Abstract
Paraphrasing uses different words, sentence structures, or expressions to convey similar semantics. It is an effective training data augmentation method to improve low-resource Natural Language Processing (NLP) tasks. Existing studies normally leverage parallel corpora to construct parabanks, regarding the Machine Translation (MT) results of source sentences as the paraphrases of the corresponding target sentences. As MT models are usually trained on the same parallel corpus, translation of the training set may suffer from overfitting, which leads to less diverse paraphrases. Training paraphrasers on the parabank generated via MT may also suffer from the information loss issue, as the parabank is derived from the parallel corpora, and the knowledge inside the parabank is a subset of that inside the parallel corpora. In this paper, we train bidirectional Multilingual Neural Machine Translation (MNMT) on the bi-directional bilingual parallel corpus, and use the MNMT model directly as a paraphrasing model by asking it to generate "translations" of the input language. As some source tokens also appear in the translation in the parallel corpus, we introduce "copy"/"not-copy" tags to indicate the existence/non-existence of source tokens in the target translation during training, and use the "not-copy" tag to encourage paraphrasing during inference. Manual and automatic evaluation results show that our ParaMNMT method can generate paraphrases of higher semantic consistency, literal fluency and sentential diversity compared to existing parabanks and LLMs. Our data augmentation experiments verify the effectiveness of ParaM-NMT on improving low-resource NLP tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on19
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Prompting Large Language Model for Machine Translation: A Case StudyBiao Zhang, Barry Haddow, Alexandra BirchICML 2023 · 402 citations
- Improving Massively Multilingual Neural Machine Translation and Zero-Shot TranslationBiao Zhang, Philip Williams, Ivan Titov, Rico SennrichACL 2020 · 213 citations
- Increasing Diversity While Maintaining Accuracy: Text Data Generation with Large Language Models and Human InterventionsJohn Joon Young Chung, Ece Kamar, Saleema AmershiACL 2023 · 55 citations
Related papers
- LAMPAT: Low-Rank Adaption for Multilingual Paraphrasing Using Adversarial TrainingKhoi M. Le, Trinh Pham, Tho Quan, Anh Tuan LuuAAAI 2024 · 12 citations
- ParaAMR: A Large-Scale Syntactically Diverse Paraphrase Dataset by AMR Back-TranslationKuan-Hao Huang, Varun Iyer, I-Hung Hsu, Anoop Kumar et al.ACL 2023 · 3 citations
- Rethinking Data Augmentation for Low-Resource Neural Machine Translation: A Multi-Task Learning ApproachVíctor M. Sánchez-Cartagena, Miquel Esplà-Gomis, Juan Antonio Pérez-Ortiz, Felipe Sánchez-MartínezEMNLP 2021 · 21 citations
- Selecting Backtranslated Data from Multiple Sources for Improved Neural Machine TranslationXabier Soto, Dimitar Sht. Shterionov, Alberto Poncelas, Andy WayACL 2020 · 1 citation
- Learning to Generalize to More: Continuous Semantic Augmentation for Neural Machine TranslationXiangpeng Wei, Heng Yu, Yue Hu, Rongxiang Weng et al.ACL 2022 · 26 citations
