SeqMix: Augmenting Active Sequence Labeling via Sequence Mixup
Rongzhi Zhang, Yue Yu, Chao Zhang
Abstract
Active learning is an important technique for low-resource sequence labeling tasks. However, current active sequence labeling methods use the queried samples alone in each iteration, which is an inefficient way of leveraging human annotations. We propose a simple but effective data augmentation method to improve label efficiency of active sequence labeling. Our method, SeqMix, simply augments the queried samples by generating extra labeled sequences in each iteration. The key difficulty is to generate plausible sequences along with token-level labels. In SeqMix, we address this challenge by performing mixup for both sequences and token-level labels of the queried samples. Furthermore, we design a discriminator during sequence mixup, which judges whether the generated sequences are plausible or not. Our experiments on Named Entity Recognition and Event Detection tasks show that SeqMix can improve the standard active sequence labeling method by 2.27%-3.75% in terms of F 1 scores. The code and data for SeqMix can be found at https://github. com/rz-zhang/SeqMix .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ca5f404e-58cb-4ba6-9569-d530ed709e8aCited by top-tier papers17
- C-Mixup: Improving Generalization in RegressionHuaxiu Yao, Yiping Wang, Linjun Zhang, James Y. Zou et al.NeurIPS 2022 · 106 citations
- A Survey of Active Learning for Natural Language ProcessingZhisong Zhang, Emma Strubell, Eduard H. HovyEMNLP 2022 · 60 citations
- On the Calibration of Pre-trained Language Models using Mixup Guided by Area Under the Margin and SaliencySeoyeon Park, Cornelia CarageaACL 2022 · 44 citations
- Data Augmentation for Cross-Domain Named Entity RecognitionShuguang Chen, Gustavo Aguilar, Leonardo Neves, Thamar SolorioEMNLP 2021 · 39 citations
- ALP: Data Augmentation Using Lexicalized PCFGs for Few-Shot Text ClassificationHazel H. Kim, Daecheol Woo, Seong Joon Oh, Jeong-Won Cha et al.AAAI 2022 · 33 citations
Builds on7
- Unsupervised Data Augmentation for Consistency TrainingQizhe Xie, Zihang Dai, Eduard H. Hovy, Thang Luong et al.NeurIPS 2020 · 2,774 citations
- Do Not Have Enough Data? Deep Learning to the Rescue!Ateret Anaby-Tavor, Boaz Carmeli, Esther Goldbraich, Amir Kantor et al.AAAI 2020 · 398 citations
- MixText: Linguistically-Informed Interpolation of Hidden Space for Semi-Supervised Text ClassificationJiaao Chen, Zichao Yang, Diyi YangACL 2020 · 340 citations
- BOND: BERT-Assisted Open-Domain Named Entity Recognition with Distant SupervisionChen Liang, Yue Yu, Haoming Jiang, Siawpeng Er et al.KDD 2020 · 118 citations
- SeqVAT: Virtual Adversarial Training for Semi-Supervised Sequence LabelingLuoxin Chen, Weitong Ruan, Xinyue Liu, Jianhua LuACL 2020 · 118 citations
Related papers
- EASAL: Entity-Aware Subsequence-Based Active Learning for Named Entity RecognitionYang Liu, Jinpeng Hu, Zhihong Chen, Xiang Wan et al.AAAI 2023 · 2 citations
- Robust and Informative Text Augmentation (RITA) via Constrained Worst-Case Transformations for Low-Resource Named Entity RecognitionHyunwoo Sohn, Baekkwan ParkKDD 2022 · 3 citations
- Adversarial Word Dilution as Text Data Augmentation in Low-Resource RegimeJunfan Chen, Richong Zhang, Zheyan Luo, Chunming Hu et al.AAAI 2023 · 7 citations
- Meta Self-training for Few-shot Neural Sequence LabelingYaqing Wang, Subhabrata Mukherjee, Haoda Chu, Yuancheng Tu et al.KDD 2021 · 56 citations
- NAT: Noise-Aware Training for Robust Neural Sequence LabelingMarcin Namysl, Sven Behnke, Joachim KöhlerACL 2020
