Crowdsourcing Learning as Domain Adaptation: A Case Study on Named Entity Recognition
Xin Zhang, Guangwei Xu, Yueheng Sun, Meishan Zhang, Pengjun Xie
Abstract
Crowdsourcing is regarded as one prospective solution for effective supervised learning, aiming to build large-scale annotated training data by crowd workers. Previous studies focus on reducing the influences from the noises of the crowdsourced annotations for supervised models. We take a different point in this work, regarding all crowdsourced annotations as goldstandard with respect to the individual annotators. In this way, we find that crowdsourcing could be highly similar to domain adaptation, and then the recent advances of cross-domain methods can be almost directly applied to crowdsourcing. Here we take named entity recognition (NER) as a study case, suggesting an annotator-aware representation learning model that inspired by the domain adaptation methods which attempt to capture effective domain-aware features. We investigate both unsupervised and supervised crowdsourcing learning, assuming that no or only smallscale expert annotations are available. Experimental results on a benchmark crowdsourced NER dataset show that our method is highly effective, leading to a new state-of-the-art performance. In addition, under the supervised setting, we can achieve impressive performance gains with only a very small scale of expert annotations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c020d59d-8a8a-4bb4-b5d5-597739b2f3a3Cited by top-tier papers7
- PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning OptimizationYidong Wang, Zhuohao Yu, Wenjin Yao, Zhengran Zeng et al.ICLR 2024 · 368 citations
- Unsupervised Entity Alignment for Temporal Knowledge GraphsXiaoze Liu, Junyang Wu, Tianyi Li, Lu Chen et al.WWW 2023 · 56 citations
- Character-level White-Box Adversarial Attacks against Transformers via Attachable Subwords SubstitutionAiwei Liu, Honghai Yu, Xuming Hu, Shu'ang Li et al.EMNLP 2022 · 20 citations
- Identifying Chinese Opinion Expressions with Extremely-Noisy Crowdsourcing AnnotationsXin Zhang, Guangwei Xu, Yueheng Sun, Meishan Zhang et al.ACL 2022 · 12 citations
- Query-based Instance Discrimination Network for Relational Triple ExtractionZeqi Tan, Yongliang Shen, Xuming Hu, Wenqi Zhang et al.EMNLP 2022 · 10 citations
Builds on4
- Multi-Cell Compositional LSTM for NER Domain AdaptationChen Jia, Yue ZhangACL 2020 · 62 citations
- Robust Named Entity Recognition with Truecasing PretrainingStephen Mayhew, Nitish Gupta, Dan RothAAAI 2020 · 43 citations
- Multi-Domain Named Entity Recognition with Genre-Aware and Agnostic InferenceJing Wang, Mayank Kulkarni, Daniel Preotiuc-PietroACL 2020 · 31 citations
- UDapter: Language Adaptation for Truly Universal Dependency ParsingAhmet Üstün, Arianna Bisazza, Gosse Bouma, Gertjan van NoordEMNLP 2020 · 10 citations
Related papers
- Zero-Resource Cross-Lingual Named Entity RecognitionM. Saiful Bari, Shafiq R. Joty, Prathyusha JwalapuramAAAI 2020 · 55 citations
- MetaNER: Named Entity Recognition with Meta-LearningJing Li, Shuo Shang, Ling ShaoWWW 2020 · 56 citations
- Named Entity Recognition without Labelled Data: A Weak Supervision ApproachPierre Lison, Jeremy Barnes, Aliaksandr Hubin, Samia TouilebACL 2020 · 12 citations
- Few-Shot Named Entity Recognition: An Empirical Baseline StudyJiaxin Huang, Chunyuan Li, Krishan Subudhi, Damien Jose et al.EMNLP 2021 · 97 citations
- Single-/Multi-Source Cross-Lingual NER via Teacher-Student Learning on Unlabeled Data in Target LanguageQianhui Wu, Zijia Lin, Börje Karlsson, Jianguang Lou et al.ACL 2020 · 59 citations
