Improving Paraphrase Detection with the Adversarial Paraphrasing Task
Animesh Nighojkar, John Licato
Abstract
If two sentences have the same meaning, it should follow that they are equivalent in their inferential properties, i.e., each sentence should textually entail the other. However, many paraphrase datasets currently in widespread use rely on a sense of paraphrase based on word overlap and syntax. Can we teach them instead to identify paraphrases in a way that draws on the inferential properties of the sentences, and is not over-reliant on lexical and syntactic similarities of a sentence pair? We apply the adversarial paradigm to this question, and introduce a new adversarial method of dataset creation for paraphrase identification: the Adversarial Paraphrasing Task (APT), which asks participants to generate semantically equivalent (in the sense of mutually implicative) but lexically and syntactically disparate paraphrases. These sentence pairs can then be used both to test paraphrase identification models (which get barely random accuracy) and then improve their performance. To accelerate dataset generation, we explore automation of APT using T5, and show that the resulting dataset also improves accuracy. We discuss implications for paraphrase detection and release our dataset in the hope of making paraphrase detection models better able to detect sentence-level meaning equivalence.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 77669013-5ae7-46c9-a7f5-0b2a762e88eaCited by top-tier papers7
- FACTIFY-5WQA: 5W Aspect-based Fact Verification through Question AnsweringAnku Rani, S. M. Towhidul Islam Tonmoy, Dwip Dalal, Shreya Gautam et al.ACL 2023 · 16 citations
- Self-Supervised Query Reformulation for Code SearchYuetian Mao, Chengcheng Wan, Yuze Jiang, Xiaodong GuFSE 2023 · 14 citations
- Labels Need Prompts Too: Mask Matching for Natural Language Understanding TasksBo Li, Wei Ye, Quansen Wang, Wen Zhao et al.AAAI 2024 · 4 citations
- FACTIFY3M: A benchmark for multimodal fact verification with explainability through 5W Question-AnsweringMegha Chakraborty, Khushbu Pahwa, Anku Rani, Shreyas Chatterjee et al.EMNLP 2023 · 3 citations
- Modeling Information Change in Science Communication with Semantically Matched ParaphrasesDustin Wright, Jiaxin Pei, David Jurgens, Isabelle AugensteinEMNLP 2022 · 2 citations
Builds on5
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- Adversarial NLI: A New Benchmark for Natural Language UnderstandingYixin Nie, Adina Williams, Emily Dinan, Mohit Bansal et al.ACL 2020 · 602 citations
- Adversarial Filters of Dataset BiasesRonan Le Bras, Swabha Swayamdipta, Chandra Bhagavatula, Rowan Zellers et al.ICML 2020 · 242 citations
- BLEURT: Learning Robust Metrics for Text GenerationThibault Sellam, Dipanjan Das, Ankur P. ParikhACL 2020 · 40 citations
- How to Ask Good Questions? Try to Leverage ParaphrasesXin Jia, Wenjie Zhou, Xu Sun, Yunfang WuACL 2020 · 29 citations
Related papers
- Entailment Relation Aware Paraphrase GenerationAbhilasha Sancheti, Balaji Vasan Srinivasan, Rachel RudingerAAAI 2022 · 5 citations
- ParaAMR: A Large-Scale Syntactically Diverse Paraphrase Dataset by AMR Back-TranslationKuan-Hao Huang, Varun Iyer, I-Hung Hsu, Anoop Kumar et al.ACL 2023 · 3 citations
- LAMPAT: Low-Rank Adaption for Multilingual Paraphrasing Using Adversarial TrainingKhoi M. Le, Trinh Pham, Tho Quan, Anh Tuan LuuAAAI 2024 · 12 citations
- Adversarial Semantic CollisionsCongzheng Song, Alexander M. Rush, Vitaly ShmatikovEMNLP 2020 · 31 citations
- PARADE: A New Dataset for Paraphrase Identification Requiring Computer Science Domain KnowledgeYun He, Zhuoer Wang, Yin Zhang, Ruihong Huang et al.EMNLP 2020 · 14 citations
