Low Resource Sequence Tagging with Weak Labels
Edwin Simpson, Jonas Pfeiffer, Iryna Gurevych
摘要
Current methods for sequence tagging depend on large quantities of domain-specific training data, limiting their use in new, user-defined tasks with few or no annotations. While crowdsourcing can be a cheap source of labels, it often introduces errors that degrade the performance of models trained on such crowdsourced data. Another solution is to use transfer learning to tackle low resource sequence labelling, but current approaches rely heavily on similar high resource datasets in different languages. In this paper, we propose a domain adaptation method using Bayesian sequence combination to exploit pre-trained models and unreliable crowdsourced data that does not require high resource data in a different language. Our method boosts performance by learning the relationship between each labeller and the target task and trains a sequence labeller on the target domain with little or no gold-standard data. We apply our approach to labelling diagnostic classes in medical and educational case studies, showing that the model achieves strong performance though zero-shot transfer learning and is more effective than alternative ensemble methods. Using NER and information extraction tasks, we show how our approach can train a model directly from crowdsourced labels, outperforming pipeline approaches that first aggregate the crowdsourced data, then train on the aggregated labels.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Learning to Contextually Aggregate Multi-Source Supervision for Sequence LabelingOuyu Lan, Xiao Huang, Bill Yuchen Lin, He Jiang 等ACL 2020 · 被引用 33 次
- PDALN: Progressive Domain Adaptation over a Pre-trained Model for Low-Resource Cross-Domain Named Entity RecognitionTao Zhang, Congying Xia, Philip S. Yu, Zhiwei Liu 等EMNLP 2021 · 被引用 22 次
- Agreeing to Disagree: Annotating Offensive Language Datasets with Annotators' DisagreementElisa Leonardelli, Stefano Menini, Alessio Palmero Aprosio, Marco Guerini 等EMNLP 2021 · 被引用 2 次
相关 Paper
- Named Entity Recognition without Labelled Data: A Weak Supervision ApproachPierre Lison, Jeremy Barnes, Aliaksandr Hubin, Samia TouilebACL 2020 · 被引用 12 次
- Crowdsourcing Learning as Domain Adaptation: A Case Study on Named Entity RecognitionXin Zhang, Guangwei Xu, Yueheng Sun, Meishan Zhang 等ACL 2021
- Semi-Supervised Knowledge Amalgamation for Sequence ClassificationJidapa Thadajarassiri, Thomas Hartvigsen, Xiangnan Kong, Elke A. RundensteinerAAAI 2021 · 被引用 14 次
- Argument Mining in Data Scarce Settings: Cross-lingual Transfer and Few-shot TechniquesAnar Yeginbergen, Maite Oronoz, Rodrigo AgerriACL 2024
- UXLA: A Robust Unsupervised Data Augmentation Framework for Zero-Resource Cross-Lingual NLPM. Saiful Bari, Tasnim Mohiuddin, Shafiq R. JotyACL 2021
