Sparse Conditional Hidden Markov Model for Weakly Supervised Named Entity Recognition
Yinghao Li, Le Song, Chao Zhang
摘要
Weakly supervised named entity recognition methods train label models to aggregate the token annotations of multiple noisy labeling functions (LFs) without seeing any manually annotated labels. To work well, the label model needs to contextually identify and emphasize well-performed LFs while down-weighting the under-performers. However, evaluating the LFs is challenging due to the lack of ground truths. To address this issue, we propose the sparse conditional hidden Markov model (Sparse-CHMM). Instead of predicting the entire emission matrix as other HMM-based methods, Sparse-CHMM focuses on estimating its diagonal elements, which are considered as the reliability scores of the LFs. The sparse scores are then expanded to the full-fledged emission matrix with pre-defined expansion functions. We also augment the emission with weighted XOR scores, which track the probabilities of an LF observing incorrect entities. Sparse-CHMM is optimized through unsupervised learning with a three-stage training pipeline that reduces the training difficulty and prevents the model from falling into local optima. Compared with the baselines in the Wrench benchmark, Sparse-CHMM achieves a 3.01 average F1 score improvement on five comprehensive datasets. Experiments show that each component of Sparse-CHMM is effective, and the estimated LF reliabilities strongly correlate with true LF F1 scores.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper6
- Weakly Supervised Sequence Tagging from Noisy RulesEsteban Safranchik, Shiying Luo, Stephen H. BachAAAI 2020 · 被引用 90 次
- Learning to Contextually Aggregate Multi-Source Supervision for Sequence LabelingOuyu Lan, Xiao Huang, Bill Yuchen Lin, He Jiang 等ACL 2020 · 被引用 33 次
- Named Entity Recognition without Labelled Data: A Weak Supervision ApproachPierre Lison, Jeremy Barnes, Aliaksandr Hubin, Samia TouilebACL 2020 · 被引用 12 次
- Interactive Weak Supervision: Learning Useful Heuristics for Data LabelingBenedikt Boecking, Willie Neiswanger, Eric P. Xing, Artur DubrawskiICLR 2021 · 被引用 8 次
- DrNAS: Dirichlet Neural Architecture SearchXiangning Chen, Ruochen Wang, Minhao Cheng, Xiaocheng Tang 等ICLR 2021 · 被引用 7 次
相关 Paper
- BERTifying the Hidden Markov Model for Multi-Source Weakly Supervised Named Entity RecognitionYinghao Li, Pranav Shetty, Lucas Liu, Chao Zhang 等ACL 2021
- Addressing NER Annotation Noises with Uncertainty-Guided Tree-Structured CRFsJian Liu, Weichang Liu, Yufeng Chen, Jinan Xu 等EMNLP 2023 · 被引用 3 次
- Named Entity Recognition Only from Word EmbeddingsYing Luo, Hai Zhao, Junlang ZhanEMNLP 2020 · 被引用 22 次
- Evaluating multiple models using labeled and unlabeled dataDivya Shanmugam, Shuvom Sadhuka, Manish Raghavan, John V. Guttag 等NeurIPS 2025 · 被引用 9 次
- Distantly-Supervised Named Entity Recognition with Noise-Robust Learning and Language Model Augmented Self-TrainingYu Meng, Yunyi Zhang, Jiaxin Huang, Xuan Wang 等EMNLP 2021 · 被引用 50 次
