Towards Better Characterization of Paraphrases
Timothy Liu, De Wen Soh
Abstract
To effectively characterize the nature of paraphrase pairs without expert human annotation, we proposes two new metrics: word position deviation (WPD) and lexical deviation (LD). WPD measures the degree of structural alteration, while LD measures the difference in vocabulary used. We apply these metrics to better understand the commonly-used MRPC dataset and study how it differs from PAWS, another paraphrase identification dataset. We also perform a detailed study on MRPC and propose improvements to the dataset, showing that it improves generalizability of models trained on the dataset. Lastly, we apply our metrics to filter the output of a paraphrase generation model and show how it can be used to generate specific forms of paraphrases for data augmentation or robustness testing of NLP models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 482a32dc-d17c-489f-908f-9197392b829aCited by top-tier papers5
- Improving Large-scale Paraphrase Acquisition and GenerationYao Dou, Chao Jiang, Wei XuEMNLP 2022 · 11 citations
- Paraphrase Types for Generation and DetectionJan Philip Wahle, Bela Gipp, Terry RuasEMNLP 2023 · 7 citations
- Paraphrase Types Elicit Prompt Engineering CapabilitiesJan Philip Wahle, Terry Ruas, Yang Xu, Bela GippEMNLP 2024 · 4 citations
- OssCSE: Overcoming Surface Structure Bias in Contrastive Learning for Unsupervised Sentence EmbeddingZhan Shi, Guoyin Wang, Ke Bai, Jiwei Li et al.EMNLP 2023 · 3 citations
- FCMR: Robust Evaluation of Financial Cross-Modal Multi-Hop ReasoningSeunghee Kim, Changhyeon Kim, Taeuk KimACL 2025
Builds on2
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 3,729 citations
- Competency Problems: On Finding and Removing Artifacts in Language DataMatt Gardner, William Merrill, Jesse Dodge, Matthew E. Peters et al.EMNLP 2021 · 72 citations
Related papers
- On the Evaluation Metrics for Paraphrase GenerationLingfeng Shen, Lemao Liu, Haiyun Jiang, Shuming ShiEMNLP 2022 · 29 citations
- ParaTag: A Dataset of Paraphrase Tagging for Fine-Grained Labels, NLG Evaluation, and Data AugmentationShuohang Wang, Ruochen Xu, Yang Liu, Chenguang Zhu et al.EMNLP 2022 · 3 citations
- ParaAMR: A Large-Scale Syntactically Diverse Paraphrase Dataset by AMR Back-TranslationKuan-Hao Huang, Varun Iyer, I-Hung Hsu, Anoop Kumar et al.ACL 2023 · 3 citations
- GAPX: Generalized Autoregressive Paraphrase-Identification XYifei Zhou, Renyu Li, Hayden Housen, Ser Nam LimNeurIPS 2022 · 1 citation
- PARADE: A New Dataset for Paraphrase Identification Requiring Computer Science Domain KnowledgeYun He, Zhuoer Wang, Yin Zhang, Ruihong Huang et al.EMNLP 2020 · 14 citations
