ParaAMR: A Large-Scale Syntactically Diverse Paraphrase Dataset by AMR Back-Translation
Kuan-Hao Huang, Varun Iyer, I-Hung Hsu, Anoop Kumar, Kai-Wei Chang, Aram Galstyan
摘要
Paraphrase generation is a long-standing task in natural language processing (NLP). Supervised paraphrase generation models, which rely on human-annotated paraphrase pairs, are costinefficient and hard to scale up. On the other hand, automatically annotated paraphrase pairs (e.g., by machine back-translation), usually suffer from the lack of syntactic diversity -the generated paraphrase sentences are very similar to the source sentences in terms of syntax. In this work, we present PARAAMR, a large-scale syntactically diverse paraphrase dataset created by abstract meaning representation backtranslation. Our quantitative analysis, qualitative examples, and human evaluation demonstrate that the paraphrases of PARAAMR are syntactically more diverse compared to existing large-scale paraphrase datasets while preserving good semantic similarity. In addition, we show that PARAAMR can be used to improve on three NLP tasks: learning sentence embeddings, syntactically controlled paraphrase generation, and data augmentation for few-shot learning. Our results thus showcase the potential of PARAAMR for improving various NLP applications.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- A Survey of AMR ApplicationsShira Wein, Juri OpitzEMNLP 2024 · 被引用 7 次
- How Far Can We Extract Diverse Perspectives from Large Language Models?Shirley Anugrah Hayati, Minhwa Lee, Dheeraj Rajagopal, Dongyeop KangEMNLP 2024 · 被引用 6 次
- Data Drives Unstable Hierarchical Generalization in LMsTian Qin, Naomi Saphra, David Alvarez-MelisEMNLP 2025
- Paraphrasing as Zero-shot Translation with Feature-guided Diversity EnhancementZiyue Yan, Hongying Zan, Xinglin Lyu, Hongfei XuACL 2026
它引用的顶会 Paper9
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 被引用 2,496 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- One SPRING to Rule Them Both: Symmetric AMR Semantic Parsing and Generation without a Complex PipelineMichele Bevilacqua, Rexhina Blloshmi, Roberto NavigliAAAI 2021 · 被引用 197 次
- AMR Parsing via Graph-Sequence Iterative InferenceDeng Cai, Wai LamACL 2020 · 被引用 83 次
- Bridging the Structural Gap Between Encoding and Decoding for Data-To-Text GenerationChao Zhao, Marilyn A. Walker, Snigdha ChaturvediACL 2020 · 被引用 82 次
相关 Paper
- LAMPAT: Low-Rank Adaption for Multilingual Paraphrasing Using Adversarial TrainingKhoi M. Le, Trinh Pham, Tho Quan, Anh Tuan LuuAAAI 2024 · 被引用 12 次
- ABEX: Data Augmentation for Low-Resource NLU via Expanding Abstract DescriptionsSreyan Ghosh, Utkarsh Tyagi, Sonal Kumar, Chandra Kiran Reddy Evuru 等ACL 2024 · 被引用 3 次
- Revisiting Pivot-Based Paraphrase Generation: Language Is Not the Only Optional PivotYitao Cai, Yue Cao, Xiaojun WanEMNLP 2021 · 被引用 6 次
- Unifying Discrete and Continuous Representations for Unsupervised Paraphrase GenerationMingfeng Xue, Dayiheng Liu, Wenqiang Lei, Jie Fu 等EMNLP 2023 · 被引用 2 次
- Retrofitting Multilingual Sentence Embeddings with Abstract Meaning RepresentationDeng Cai, Xin Li, Jackie Chun-Sing Ho, Lidong Bing 等EMNLP 2022 · 被引用 4 次
