ParaTag: A Dataset of Paraphrase Tagging for Fine-Grained Labels, NLG Evaluation, and Data Augmentation
Shuohang Wang, Ruochen Xu, Yang Liu, Chenguang Zhu, Michael Zeng
Abstract
Paraphrase identification has been formulated as a binary classification task to decide whether two sentences hold a paraphrase relationship. Existing paraphrase datasets only annotate a binary label for each sentence pair. However, after a systematical analysis of existing paraphrase datasets, we found that the degree of paraphrase cannot be well characterized by a single binary label. And the criteria of paraphrase are not even consistent within the same dataset. We hypothesize that such issues would limit the effectiveness of paraphrase models trained on these data. To this end, we propose a novel fine-grained paraphrase annotation schema that labels the minimum spans of tokens in a sentence that don't have the corresponding paraphrases in the other sentence. Under this setting, we frame paraphrasing as a sequence tagging task. We collect 30k sentence pairs in English with the new annotation schema, resulting in the ParaTag dataset. In addition to reporting baseline results on ParaTag using state-of-art language models, we show that ParaTag is especially useful for training an automatic scorer for language generation evaluation. Finally, we train a paraphrase generation model from ParaTag and achieve better data augmentation performance on the GLUE benchmark than other public paraphrasing datasets. 1 * Equal contribution. 1 Data and code: https://github.com/microsoft/ParaTag M-1 However, commercial use of the 2.6.0 kernel is still months off for most customers. M-2 Commercial releases of the 2.6 kernel by major Linux distributors still remain months away.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Paraphrasing as Zero-shot Translation with Feature-guided Diversity EnhancementZiyue Yan, Hongying Zan, Xinglin Lyu, Hongfei XuACL 2026
- What's Mine becomes Yours: Defining, Annotating and Detecting Context-Dependent Paraphrases in News Interview DialogsAnna Wegmann, Tijs A. van den Broek, Dong NguyenEMNLP 2024
Builds on10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- BARTScore: Evaluating Generated Text as Text GenerationWeizhe Yuan, Graham Neubig, Pengfei LiuNeurIPS 2021 · 1,143 citations
- DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding SharingPengcheng He, Jianfeng Gao, Weizhu ChenICLR 2023 · 394 citations
Related papers
- Improving Large-scale Paraphrase Acquisition and GenerationYao Dou, Chao Jiang, Wei XuEMNLP 2022 · 11 citations
- Paraphrase Types for Generation and DetectionJan Philip Wahle, Bela Gipp, Terry RuasEMNLP 2023 · 7 citations
- Towards Better Characterization of ParaphrasesTimothy Liu, De Wen SohACL 2022 · 9 citations
- PARADE: A New Dataset for Paraphrase Identification Requiring Computer Science Domain KnowledgeYun He, Zhuoer Wang, Yin Zhang, Ruihong Huang et al.EMNLP 2020 · 14 citations
- Learning to Selectively Learn for Weakly-supervised Paraphrase GenerationKaize Ding, Dingcheng Li, Alexander Hanbo Li, Xing Fan et al.EMNLP 2021 · 4 citations
