SeqMix: Augmenting Active Sequence Labeling via Sequence Mixup
Rongzhi Zhang, Yue Yu, Chao Zhang
摘要
Active learning is an important technique for low-resource sequence labeling tasks. However, current active sequence labeling methods use the queried samples alone in each iteration, which is an inefficient way of leveraging human annotations. We propose a simple but effective data augmentation method to improve label efficiency of active sequence labeling. Our method, SeqMix, simply augments the queried samples by generating extra labeled sequences in each iteration. The key difficulty is to generate plausible sequences along with token-level labels. In SeqMix, we address this challenge by performing mixup for both sequences and token-level labels of the queried samples. Furthermore, we design a discriminator during sequence mixup, which judges whether the generated sequences are plausible or not. Our experiments on Named Entity Recognition and Event Detection tasks show that SeqMix can improve the standard active sequence labeling method by 2.27%-3.75% in terms of F 1 scores. The code and data for SeqMix can be found at https://github. com/rz-zhang/SeqMix .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- C-Mixup: Improving Generalization in RegressionHuaxiu Yao, Yiping Wang, Linjun Zhang, James Y. Zou 等NeurIPS 2022 · 被引用 106 次
- A Survey of Active Learning for Natural Language ProcessingZhisong Zhang, Emma Strubell, Eduard H. HovyEMNLP 2022 · 被引用 60 次
- On the Calibration of Pre-trained Language Models using Mixup Guided by Area Under the Margin and SaliencySeoyeon Park, Cornelia CarageaACL 2022 · 被引用 44 次
- Data Augmentation for Cross-Domain Named Entity RecognitionShuguang Chen, Gustavo Aguilar, Leonardo Neves, Thamar SolorioEMNLP 2021 · 被引用 39 次
- ALP: Data Augmentation Using Lexicalized PCFGs for Few-Shot Text ClassificationHazel H. Kim, Daecheol Woo, Seong Joon Oh, Jeong-Won Cha 等AAAI 2022 · 被引用 33 次
它引用的顶会 Paper7
- Unsupervised Data Augmentation for Consistency TrainingQizhe Xie, Zihang Dai, Eduard H. Hovy, Thang Luong 等NeurIPS 2020 · 被引用 2,774 次
- Do Not Have Enough Data? Deep Learning to the Rescue!Ateret Anaby-Tavor, Boaz Carmeli, Esther Goldbraich, Amir Kantor 等AAAI 2020 · 被引用 398 次
- MixText: Linguistically-Informed Interpolation of Hidden Space for Semi-Supervised Text ClassificationJiaao Chen, Zichao Yang, Diyi YangACL 2020 · 被引用 340 次
- BOND: BERT-Assisted Open-Domain Named Entity Recognition with Distant SupervisionChen Liang, Yue Yu, Haoming Jiang, Siawpeng Er 等KDD 2020 · 被引用 118 次
- SeqVAT: Virtual Adversarial Training for Semi-Supervised Sequence LabelingLuoxin Chen, Weitong Ruan, Xinyue Liu, Jianhua LuACL 2020 · 被引用 118 次
相关 Paper
- EASAL: Entity-Aware Subsequence-Based Active Learning for Named Entity RecognitionYang Liu, Jinpeng Hu, Zhihong Chen, Xiang Wan 等AAAI 2023 · 被引用 2 次
- Robust and Informative Text Augmentation (RITA) via Constrained Worst-Case Transformations for Low-Resource Named Entity RecognitionHyunwoo Sohn, Baekkwan ParkKDD 2022 · 被引用 3 次
- Adversarial Word Dilution as Text Data Augmentation in Low-Resource RegimeJunfan Chen, Richong Zhang, Zheyan Luo, Chunming Hu 等AAAI 2023 · 被引用 7 次
- Meta Self-training for Few-shot Neural Sequence LabelingYaqing Wang, Subhabrata Mukherjee, Haoda Chu, Yuancheng Tu 等KDD 2021 · 被引用 56 次
- NAT: Noise-Aware Training for Robust Neural Sequence LabelingMarcin Namysl, Sven Behnke, Joachim KöhlerACL 2020
