PARADE: A New Dataset for Paraphrase Identification Requiring Computer Science Domain Knowledge
Yun He, Zhuoer Wang, Yin Zhang, Ruihong Huang, James Caverlee
Abstract
We present a new benchmark dataset called PARADE for paraphrase identification that requires specialized domain knowledge. PA-RADE contains paraphrases that overlap very little at the lexical and syntactic level but are semantically equivalent based on computer science domain knowledge, as well as nonparaphrases that overlap greatly at the lexical and syntactic level but are not semantically equivalent based on this domain knowledge. Experiments show that both state-of-the-art neural models and non-expert human annotators have poor performance on PARADE. For example, BERT after fine-tuning achieves an F1 score of 0.709, which is much lower than its performance on other paraphrase identification datasets. PARADE can serve as a resource for researchers interested in testing models that incorporate domain knowledge. We make our data and code freely available. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f4790f7b-a208-4cb2-9a04-4ed017e1554fBuilds on1
Related papers
- Improving Paraphrase Detection with the Adversarial Paraphrasing TaskAnimesh Nighojkar, John LicatoACL 2021
- Improving Large-scale Paraphrase Acquisition and GenerationYao Dou, Chao Jiang, Wei XuEMNLP 2022 · 11 citations
- ParaTag: A Dataset of Paraphrase Tagging for Fine-Grained Labels, NLG Evaluation, and Data AugmentationShuohang Wang, Ruochen Xu, Yang Liu, Chenguang Zhu et al.EMNLP 2022 · 3 citations
- Towards Better Characterization of ParaphrasesTimothy Liu, De Wen SohACL 2022 · 9 citations
- Unsupervised Paraphrasing under Syntax KnowledgeTianyuan Liu, Yuqing Sun, Jiaqi Wu, Xi Xu et al.AAAI 2023 · 3 citations
