PeerDA: Data Augmentation via Modeling Peer Relation for Span Identification Tasks
Weiwen Xu, Xin Li, Yang Deng, Wai Lam, Lidong Bing
摘要
Span identification aims at identifying specific text spans from text input and classifying them into pre-defined categories. Different from previous works that merely leverage the Subordinate (SUB) relation (i.e. if a span is an instance of a certain category) to train models, this paper for the first time explores the Peer (PR) relation, which indicates that two spans are instances of the same category and share similar features. Specifically, a novel Peer Data Augmentation (PeerDA) approach is proposed which employs span pairs with the PR relation as the augmentation data for training. PeerDA has two unique advantages: (1) There are a large number of PR span pairs for augmenting the training data. (2) The augmented data can prevent the trained model from over-fitting the superficial span-category mapping by pushing the model to leverage the span semantics. Experimental results on ten datasets over four diverse tasks across seven domains demonstrate the effectiveness of PeerDA. Notably, PeerDA achieves state-of-the-art results on six of them. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- ConNER: Consistency Training for Cross-lingual Named Entity RecognitionRan Zhou, Xin Li, Lidong Bing, Erik Cambria 等EMNLP 2022 · 被引用 16 次
- From Cloze to Comprehension: Retrofitting Pre-trained Masked Language Models to Pre-trained Machine ReaderWeiwen Xu, Xin Li, Wenxuan Zhang, Meng Zhou 等NeurIPS 2023 · 被引用 3 次
- Towards Robust Low-Resource Fine-Tuning with Multi-View Compressed RepresentationsLinlin Liu, Xingxuan Li, Megh Thakkar, Xin Li 等ACL 2023 · 被引用 3 次
- Exogenous and Endogenous Data Augmentation for Low-Resource Complex Named Entity RecognitionXinghua Zhang, Gaode Chen, Shiyao Cui, Jiawei Sheng 等SIGIR 2024 · 被引用 3 次
- FineReason: Evaluating and Improving LLMs' Deliberate Reasoning through Reflective Puzzle SolvingGuizhen Chen, Weiwen Xu, Hao Zhang, Hou Pong Chan 等ACL 2025
它引用的顶会 Paper23
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 被引用 3,729 次
- Unsupervised Data Augmentation for Consistency TrainingQizhe Xie, Zihang Dai, Eduard H. Hovy, Thang Luong 等NeurIPS 2020 · 被引用 2,774 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- A Unified MRC Framework for Named Entity RecognitionXiaoya Li, Jingrong Feng, Yuxian Meng, Qinghong Han 等ACL 2020 · 被引用 617 次
相关 Paper
- Text AutoAugment: Learning Compositional Augmentation Policy for Text ClassificationShuhuai Ren, Jinchao Zhang, Lei Li, Xu Sun 等EMNLP 2021 · 被引用 22 次
- Pre-training Entity Relation Encoder with Intra-span and Inter-span InformationYijun Wang, Changzhi Sun, Yuanbin Wu, Junchi Yan 等EMNLP 2020 · 被引用 36 次
- Do Not Have Enough Data? Deep Learning to the Rescue!Ateret Anaby-Tavor, Boaz Carmeli, Esther Goldbraich, Amir Kantor 等AAAI 2020 · 被引用 398 次
- Self-Supervised Pre-training on the Target Domain for Cross-Domain Person Re-identificationJunyin Zhang, Yongxin Ge, Xinqian Gu, Boyu Hua 等ACM MM 2021 · 被引用 7 次
- Data Augmentation for Cross-Domain Named Entity RecognitionShuguang Chen, Gustavo Aguilar, Leonardo Neves, Thamar SolorioEMNLP 2021 · 被引用 39 次
