Unsupervised Paraphrasing via Deep Reinforcement Learning
A. B. Siddique, Samet Oymak, Vagelis Hristidis
Abstract
Paraphrasing is expressing the meaning of an input sentence in different wording while maintaining fluency (i.e., grammatical and syntactical correctness). Most existing work on paraphrasing use supervised models that are limited to specific domains (e.g., image captions). Such models can neither be straightforwardly transferred to other domains nor generalize well, and creating labeled training data for new domains is expensive and laborious. The need for paraphrasing across different domains and the scarcity of labeled training data in many such domains call for exploring unsupervised paraphrase generation methods. We propose Progressive Unsupervised Paraphrasing (PUP): a novel unsupervised paraphrase generation method based on deep reinforcement learning (DRL). PUP uses a variational autoencoder (trained using a non-parallel corpus) to generate a seed paraphrase that warm-starts the DRL model. Then, PUP progressively tunes the seed paraphrase guided by our novel reward function which combines semantic adequacy, language fluency, and expression diversity measures to quantify the quality of the generated paraphrases in each iteration without needing parallel sentences. Our extensive experimental evaluation shows that PUP outperforms unsupervised state-of-the-art paraphrasing techniques in terms of both automatic metrics and user studies on four real datasets. We also show that PUP outperforms domain-adapted supervised algorithms on several datasets. Our evaluation also shows that PUP achieves a great trade-off between semantic similarity and diversity of expression.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bbe8e1ec-afe9-430b-94a9-69a714f0e1efCited by top-tier papers9
- RewriteLM: An Instruction-Tuned Large Language Model for Text RewritingLei Shu, Liangchen Luo, Jayakumar Hoskere, Yun Zhu et al.AAAI 2024 · 92 citations
- On the Evaluation Metrics for Paraphrase GenerationLingfeng Shen, Lemao Liu, Haiyun Jiang, Shuming ShiEMNLP 2022 · 29 citations
- ConRPG: Paraphrase Generation using Contexts as RegularizerYuxian Meng, Xiang Ao, Qing He, Xiaofei Sun et al.EMNLP 2021 · 20 citations
- Text Revision By On-the-Fly Representation OptimizationJingjing Li, Zichao Li, Tao Ge, Irwin King et al.AAAI 2022 · 20 citations
- Unsupervised Paraphrasing with Pretrained Language ModelsTong Niu, Semih Yavuz, Yingbo Zhou, Nitish Shirish Keskar et al.EMNLP 2021 · 18 citations
Builds on1
Related papers
- Entailment Relation Aware Paraphrase GenerationAbhilasha Sancheti, Balaji Vasan Srinivasan, Rachel RudingerAAAI 2022 · 5 citations
- Unifying Discrete and Continuous Representations for Unsupervised Paraphrase GenerationMingfeng Xue, Dayiheng Liu, Wenqiang Lei, Jie Fu et al.EMNLP 2023 · 2 citations
- Generating Diverse and Descriptive Image Captions Using Visual ParaphrasesLixin Liu, Jiajun Tang, Xiaojun Wan, Zongming GuoICCV 2019 · 48 citations
- Syntactically-Informed Unsupervised Paraphrasing with Non-Parallel DataErguang Yang, Mingtong Liu, Deyi Xiong, Yujie Zhang et al.EMNLP 2021 · 6 citations
- Learning to Selectively Learn for Weakly-supervised Paraphrase GenerationKaize Ding, Dingcheng Li, Alexander Hanbo Li, Xing Fan et al.EMNLP 2021 · 4 citations
