CycleNER: An Unsupervised Training Approach for Named Entity Recognition
Andrea Iovine, Anjie Fang, Besnik Fetahu, Oleg Rokhlenko, Shervin Malmasi
摘要
Named Entity Recognition (NER) is a crucial natural language understanding task for many down-stream tasks such as question answering and retrieval. Despite significant progress in developing NER models for multiple languages and domains, scaling to emerging and/or low-resource domains still remains challenging, due to the costly nature of acquiring training data. We propose CycleNER, an unsupervised approach based on cycle-consistency training that uses two functions: (i) sentence-to-entity – S2E and (ii) entity-to-sentence – E2S, to carry out the NER task. CycleNER does not require annotations but a set of sentences with no entity labels and another independent set of entity examples. Through cycle-consistency training, the output from one function is used as input for the other (e.g. S2E → E2S) to align the representation spaces of both functions and therefore enable unsupervised training. Evaluation on several domains comparing CycleNER against supervised and unsupervised competitors shows that CycleNER achieves highly competitive performance with only a few thousand input sentences. We demonstrate competitive performance against supervised models, achieving 73% of supervised performance without any annotations on CoNLL03, while significantly outperforming unsupervised approaches.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Faithful Low-Resource Data-to-Text Generation through Cycle TrainingZhuoer Wang, Marcus D. Collins, Nikhita Vedula, Simone Filice 等ACL 2023 · 被引用 3 次
- CycleKQR: Unsupervised Bidirectional Keyword-Question RewritingAndrea Iovine, Anjie Fang, Besnik Fetahu, Jie Zhao 等EMNLP 2022 · 被引用 3 次
它引用的顶会 Paper3
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- Dice Loss for Data-imbalanced NLP TasksXiaoya Li, Xiaofei Sun, Yuxian Meng, Junjun Liang 等ACL 2020 · 被引用 575 次
- LUKE: Deep Contextualized Entity Representations with Entity-aware Self-attentionIkuya Yamada, Akari Asai, Hiroyuki Shindo, Hideaki Takeda 等EMNLP 2020 · 被引用 562 次
相关 Paper
- ConsistNER: Towards Instructive NER Demonstrations for LLMs with the Consistency of Ontology and ContextChenxiao Wu, Wenjun Ke, Peng Wang, Zhizhao Luo 等AAAI 2024 · 被引用 15 次
- Named Entity Recognition without Labelled Data: A Weak Supervision ApproachPierre Lison, Jeremy Barnes, Aliaksandr Hubin, Samia TouilebACL 2020 · 被引用 12 次
- A Unified Generative Framework for Various NER SubtasksHang Yan, Tao Gui, Junqi Dai, Qipeng Guo 等ACL 2021
- Entity Extraction in Low Resource Domains with Selective Pre-training of Large Language ModelsAniruddha Mahapatra, Sharmila Reddy Nangi, Aparna Garimella, Anandhavelu NatarajanEMNLP 2022 · 被引用 4 次
- Zero-Resource Cross-Lingual Named Entity RecognitionM. Saiful Bari, Shafiq R. Joty, Prathyusha JwalapuramAAAI 2020 · 被引用 55 次
