PeerDA: Data Augmentation via Modeling Peer Relation for Span Identification Tasks
Weiwen Xu, Xin Li, Yang Deng, Wai Lam, Lidong Bing
Abstract
Span identification aims at identifying specific text spans from text input and classifying them into pre-defined categories. Different from previous works that merely leverage the Subordinate (SUB) relation (i.e. if a span is an instance of a certain category) to train models, this paper for the first time explores the Peer (PR) relation, which indicates that two spans are instances of the same category and share similar features. Specifically, a novel Peer Data Augmentation (PeerDA) approach is proposed which employs span pairs with the PR relation as the augmentation data for training. PeerDA has two unique advantages: (1) There are a large number of PR span pairs for augmenting the training data. (2) The augmented data can prevent the trained model from over-fitting the superficial span-category mapping by pushing the model to leverage the span semantics. Experimental results on ten datasets over four diverse tasks across seven domains demonstrate the effectiveness of PeerDA. Notably, PeerDA achieves state-of-the-art results on six of them. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 91e68d0b-76ab-41b0-a351-0228e4eb9d83Cited by top-tier papers5
- ConNER: Consistency Training for Cross-lingual Named Entity RecognitionRan Zhou, Xin Li, Lidong Bing, Erik Cambria et al.EMNLP 2022 · 16 citations
- From Cloze to Comprehension: Retrofitting Pre-trained Masked Language Models to Pre-trained Machine ReaderWeiwen Xu, Xin Li, Wenxuan Zhang, Meng Zhou et al.NeurIPS 2023 · 3 citations
- Towards Robust Low-Resource Fine-Tuning with Multi-View Compressed RepresentationsLinlin Liu, Xingxuan Li, Megh Thakkar, Xin Li et al.ACL 2023 · 3 citations
- Exogenous and Endogenous Data Augmentation for Low-Resource Complex Named Entity RecognitionXinghua Zhang, Gaode Chen, Shiyao Cui, Jiawei Sheng et al.SIGIR 2024 · 3 citations
- FineReason: Evaluating and Improving LLMs' Deliberate Reasoning through Reflective Puzzle SolvingGuizhen Chen, Weiwen Xu, Hao Zhang, Hou Pong Chan et al.ACL 2025
Builds on23
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 3,729 citations
- Unsupervised Data Augmentation for Consistency TrainingQizhe Xie, Zihang Dai, Eduard H. Hovy, Thang Luong et al.NeurIPS 2020 · 2,774 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- A Unified MRC Framework for Named Entity RecognitionXiaoya Li, Jingrong Feng, Yuxian Meng, Qinghong Han et al.ACL 2020 · 617 citations
Related papers
- Text AutoAugment: Learning Compositional Augmentation Policy for Text ClassificationShuhuai Ren, Jinchao Zhang, Lei Li, Xu Sun et al.EMNLP 2021 · 22 citations
- Pre-training Entity Relation Encoder with Intra-span and Inter-span InformationYijun Wang, Changzhi Sun, Yuanbin Wu, Junchi Yan et al.EMNLP 2020 · 36 citations
- Do Not Have Enough Data? Deep Learning to the Rescue!Ateret Anaby-Tavor, Boaz Carmeli, Esther Goldbraich, Amir Kantor et al.AAAI 2020 · 398 citations
- Self-Supervised Pre-training on the Target Domain for Cross-Domain Person Re-identificationJunyin Zhang, Yongxin Ge, Xinqian Gu, Boyu Hua et al.ACM MM 2021 · 7 citations
- Data Augmentation for Cross-Domain Named Entity RecognitionShuguang Chen, Gustavo Aguilar, Leonardo Neves, Thamar SolorioEMNLP 2021 · 39 citations
