Improving Paraphrase Detection with the Adversarial Paraphrasing Task
Animesh Nighojkar, John Licato
摘要
If two sentences have the same meaning, it should follow that they are equivalent in their inferential properties, i.e., each sentence should textually entail the other. However, many paraphrase datasets currently in widespread use rely on a sense of paraphrase based on word overlap and syntax. Can we teach them instead to identify paraphrases in a way that draws on the inferential properties of the sentences, and is not over-reliant on lexical and syntactic similarities of a sentence pair? We apply the adversarial paradigm to this question, and introduce a new adversarial method of dataset creation for paraphrase identification: the Adversarial Paraphrasing Task (APT), which asks participants to generate semantically equivalent (in the sense of mutually implicative) but lexically and syntactically disparate paraphrases. These sentence pairs can then be used both to test paraphrase identification models (which get barely random accuracy) and then improve their performance. To accelerate dataset generation, we explore automation of APT using T5, and show that the resulting dataset also improves accuracy. We discuss implications for paraphrase detection and release our dataset in the hope of making paraphrase detection models better able to detect sentence-level meaning equivalence.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- FACTIFY-5WQA: 5W Aspect-based Fact Verification through Question AnsweringAnku Rani, S. M. Towhidul Islam Tonmoy, Dwip Dalal, Shreya Gautam 等ACL 2023 · 被引用 16 次
- Self-Supervised Query Reformulation for Code SearchYuetian Mao, Chengcheng Wan, Yuze Jiang, Xiaodong GuFSE 2023 · 被引用 14 次
- Labels Need Prompts Too: Mask Matching for Natural Language Understanding TasksBo Li, Wei Ye, Quansen Wang, Wen Zhao 等AAAI 2024 · 被引用 4 次
- FACTIFY3M: A benchmark for multimodal fact verification with explainability through 5W Question-AnsweringMegha Chakraborty, Khushbu Pahwa, Anku Rani, Shreyas Chatterjee 等EMNLP 2023 · 被引用 3 次
- Modeling Information Change in Science Communication with Semantically Matched ParaphrasesDustin Wright, Jiaxin Pei, David Jurgens, Isabelle AugensteinEMNLP 2022 · 被引用 2 次
它引用的顶会 Paper5
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- Adversarial NLI: A New Benchmark for Natural Language UnderstandingYixin Nie, Adina Williams, Emily Dinan, Mohit Bansal 等ACL 2020 · 被引用 602 次
- Adversarial Filters of Dataset BiasesRonan Le Bras, Swabha Swayamdipta, Chandra Bhagavatula, Rowan Zellers 等ICML 2020 · 被引用 242 次
- BLEURT: Learning Robust Metrics for Text GenerationThibault Sellam, Dipanjan Das, Ankur P. ParikhACL 2020 · 被引用 40 次
- How to Ask Good Questions? Try to Leverage ParaphrasesXin Jia, Wenjie Zhou, Xu Sun, Yunfang WuACL 2020 · 被引用 29 次
相关 Paper
- Entailment Relation Aware Paraphrase GenerationAbhilasha Sancheti, Balaji Vasan Srinivasan, Rachel RudingerAAAI 2022 · 被引用 5 次
- ParaAMR: A Large-Scale Syntactically Diverse Paraphrase Dataset by AMR Back-TranslationKuan-Hao Huang, Varun Iyer, I-Hung Hsu, Anoop Kumar 等ACL 2023 · 被引用 3 次
- LAMPAT: Low-Rank Adaption for Multilingual Paraphrasing Using Adversarial TrainingKhoi M. Le, Trinh Pham, Tho Quan, Anh Tuan LuuAAAI 2024 · 被引用 12 次
- Adversarial Semantic CollisionsCongzheng Song, Alexander M. Rush, Vitaly ShmatikovEMNLP 2020 · 被引用 31 次
- PARADE: A New Dataset for Paraphrase Identification Requiring Computer Science Domain KnowledgeYun He, Zhuoer Wang, Yin Zhang, Ruihong Huang 等EMNLP 2020 · 被引用 14 次
