Principled Paraphrase Generation with Parallel Corpora
Aitor Ormazabal, Mikel Artetxe, Aitor Soroa, Gorka Labaka, Eneko Agirre
摘要
Round-trip Machine Translation (MT) is a popular choice for paraphrase generation, which leverages readily available parallel corpora for supervision. In this paper, we formalize the implicit similarity function induced by this approach, and show that it is susceptible to non-paraphrase pairs sharing a single ambiguous translation. Based on these insights, we design an alternative similarity metric that mitigates this issue by requiring the entire translation distribution to match, and implement a relaxation of it through the Information Bottleneck method. Our approach incorporates an adversarial term into MT training in order to learn representations that encode as much information about the reference translation as possible, while keeping as little information about the input as possible. Paraphrases can be generated by decoding back to the source from this representation, without having to generate pivot translations. In addition to being more principled and efficient than round-trip MT, our approach offers an adjustable parameter to control the fidelity-diversity trade-off, and obtains better results in our experiments.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Why LLMs Hallucinate, and How to Get (Evidential) Closure: Perceptual, Intensional, and Extensional Learning for Faithful Natural Language GenerationAdam BouyamournEMNLP 2023 · 被引用 9 次
- APPLS: Evaluating Evaluation Metrics for Plain Language SummarizationYue Guo, Tal August, Gondy Leroy, Trevor Cohen 等EMNLP 2024 · 被引用 7 次
- Paraphrase Types for Generation and DetectionJan Philip Wahle, Bela Gipp, Terry RuasEMNLP 2023 · 被引用 7 次
- Unifying Discrete and Continuous Representations for Unsupervised Paraphrase GenerationMingfeng Xue, Dayiheng Liu, Wenqiang Lei, Jie Fu 等EMNLP 2023 · 被引用 2 次
它引用的顶会 Paper5
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- Unsupervised Data Augmentation for Consistency TrainingQizhe Xie, Zihang Dai, Eduard H. Hovy, Thang Luong 等NeurIPS 2020 · 被引用 2,774 次
- Translation Artifacts in Cross-lingual Transfer LearningMikel Artetxe, Gorka Labaka, Eneko AgirreEMNLP 2020 · 被引用 68 次
- Paraphrase Generation: A Survey of the State of the ArtJianing Zhou, Suma BhatEMNLP 2021 · 被引用 2 次
- Factorising Meaning and Form for Intent-Preserving ParaphrasingTom Hosking, Mirella LapataACL 2021
相关 Paper
- Revisiting Pivot-Based Paraphrase Generation: Language Is Not the Only Optional PivotYitao Cai, Yue Cao, Xiaojun WanEMNLP 2021 · 被引用 6 次
- BLEU might be Guilty but References are not InnocentMarkus Freitag, David Grangier, Isaac CaswellEMNLP 2020 · 被引用 13 次
- Paraphrasing as Zero-shot Translation with Feature-guided Diversity EnhancementZiyue Yan, Hongying Zan, Xinglin Lyu, Hongfei XuACL 2026
- Multi-Hypothesis Machine Translation EvaluationMarina Fomicheva, Lucia Specia, Francisco GuzmánACL 2020 · 被引用 13 次
- Unsupervised Paraphrasing with Pretrained Language ModelsTong Niu, Semih Yavuz, Yingbo Zhou, Nitish Shirish Keskar 等EMNLP 2021 · 被引用 18 次
