ConRPG: Paraphrase Generation using Contexts as Regularizer
Yuxian Meng, Xiang Ao, Qing He, Xiaofei Sun, Qinghong Han, Fei Wu, Chun Fan, Jiwei Li
摘要
A long-standing issue with paraphrase generation is how to obtain reliable supervision signals. In this paper, we propose an unsupervised paradigm for paraphrase generation based on the assumption that the probabilities of generating two sentences with the same meaning given the same context should be the same. Inspired by this fundamental idea, we propose a pipelined system which consists of paraphrase candidate generation based on contextual language models, candidate filtering using scoring functions, and paraphrase model training based on the selected candidates. The proposed paradigm offers merits over existing paraphrase generation methods: (1) using the context regularizer on meanings, the model is able to generate massive amounts of high-quality paraphrase pairs; and (2) using human-interpretable scoring functions to select paraphrase pairs from candidates, the proposed framework provides a channel for developers to intervene with the data generation process, leading to a more controllable model. Experimental results across different tasks and datasets demonstrate that the effectiveness of the proposed model in both supervised and unsupervised setups.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Paraphrasing evades detectors of AI-generated text, but retrieval is an effective defenseKalpesh Krishna, Yixiao Song, Marzena Karpinska, John Wieting 等NeurIPS 2023 · 被引用 657 次
- ParaLS: Lexical Substitution via Pretrained ParaphraserJipeng Qiang, Kang Liu, Yun Li, Yunhao Yuan 等ACL 2023 · 被引用 9 次
- Paraphrase Types for Generation and DetectionJan Philip Wahle, Bela Gipp, Terry RuasEMNLP 2023 · 被引用 7 次
- Reducing Sequence Length by Predicting Edit Spans with Large Language ModelsMasahiro Kaneko, Naoaki OkazakiEMNLP 2023 · 被引用 3 次
- Hierarchical Sketch Induction for Paraphrase GenerationTom Hosking, Hao Tang, Mirella LapataACL 2022
它引用的顶会 Paper5
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- Pre-training via ParaphrasingMike Lewis, Marjan Ghazvininejad, Gargi Ghosh, Armen Aghajanyan 等NeurIPS 2020 · 被引用 165 次
- Paraphrase Generation by Learning How to Edit from SamplesAmirhossein Kazemnejad, Mohammadreza Salehi, Mahdieh Soleymani BaghshahACL 2020 · 被引用 43 次
- Unsupervised Paraphrasing via Deep Reinforcement LearningA. B. Siddique, Samet Oymak, Vagelis HristidisKDD 2020 · 被引用 27 次
- Reformulating Unsupervised Style Transfer as Paraphrase GenerationKalpesh Krishna, John Wieting, Mohit IyyerEMNLP 2020 · 被引用 9 次
相关 Paper
- Unsupervised Paraphrasing with Pretrained Language ModelsTong Niu, Semih Yavuz, Yingbo Zhou, Nitish Shirish Keskar 等EMNLP 2021 · 被引用 18 次
- Learning to Selectively Learn for Weakly-supervised Paraphrase GenerationKaize Ding, Dingcheng Li, Alexander Hanbo Li, Xing Fan 等EMNLP 2021 · 被引用 4 次
- Unifying Discrete and Continuous Representations for Unsupervised Paraphrase GenerationMingfeng Xue, Dayiheng Liu, Wenqiang Lei, Jie Fu 等EMNLP 2023 · 被引用 2 次
- Improving Paraphrase Detection with the Adversarial Paraphrasing TaskAnimesh Nighojkar, John LicatoACL 2021
- Syntactically-Informed Unsupervised Paraphrasing with Non-Parallel DataErguang Yang, Mingtong Liu, Deyi Xiong, Yujie Zhang 等EMNLP 2021 · 被引用 6 次
