CrossNER: Evaluating Cross-Domain Named Entity Recognition
Zihan Liu, Yan Xu, Tiezheng Yu, Wenliang Dai, Ziwei Ji, Samuel Cahyawijaya, Andrea Madotto, Pascale Fung
Abstract
Cross-domain named entity recognition (NER) models are able to cope with the scarcity issue of NER samples in target domains. However, most of the existing NER benchmarks lack domain-specialized entity types or do not focus on a certain domain, leading to a less effective cross-domain evaluation. To address these obstacles, we introduce a cross-domain NER dataset (CrossNER), a fully-labeled collection of NER data spanning over five diverse domains with specialized entity categories for different domains. Additionally, we also provide a domain-related corpus since using it to continue pre-training language models (domain-adaptive pre-training) is effective for the domain adaptation. We then conduct comprehensive experiments to explore the effectiveness of leveraging different levels of the domain corpus and pre-training strategies to do domain-adaptive pre-training for the cross-domain task. Results show that focusing on the fractional corpus containing domain-specialized entities and utilizing a more challenging pre-training strategy in domain-adaptive pre-training are beneficial for the NER domain adaptation, and our proposed method can consistently outperform existing cross-domain NER baselines. Nevertheless, experiments also illustrate the challenge of this cross-domain NER task. We hope that our dataset and baselines will catalyze research in the NER domain adaptation area. The code and data are available at https://github.com/zliucr/CrossNER.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers29
- GoLLIE: Annotation Guidelines improve Zero-Shot Information-ExtractionOscar Sainz, Iker García-Ferrero, Rodrigo Agerri, Oier Lopez de Lacalle et al.ICLR 2024 · 168 citations
- Is GPT-3 a Good Data Annotator?Bosheng Ding, Chengwei Qin, Linlin Liu, Yew Ken Chia et al.ACL 2023 · 133 citations
- UniversalNER: Targeted Distillation from Large Language Models for Open Named Entity RecognitionWenxuan Zhou, Sheng Zhang, Yu Gu, Muhao Chen et al.ICLR 2024 · 118 citations
- Universal Information Extraction as Unified Semantic MatchingJie Lou, Yaojie Lu, Dai Dai, Wei Jia et al.AAAI 2023 · 96 citations
- Unified Dialog Model Pre-training for Task-Oriented Dialog Understanding and GenerationWanwei He, Yinpei Dai, Min Yang, Jian Sun et al.SIGIR 2022 · 41 citations
Builds on5
- Attention-Informed Mixed-Language Training for Zero-Shot Cross-Lingual Task-Oriented Dialogue SystemsZihan Liu, Genta Indra Winata, Zhaojiang Lin, Peng Xu et al.AAAI 2020 · 105 citations
- Don't Stop Pretraining: Adapt Language Models to Domains and TasksSuchin Gururangan, Ana Marasovic, Swabha Swayamdipta, Kyle Lo et al.ACL 2020 · 93 citations
- Multi-Cell Compositional LSTM for NER Domain AdaptationChen Jia, Yue ZhangACL 2020 · 62 citations
- Multi-Domain Named Entity Recognition with Genre-Aware and Agnostic InferenceJing Wang, Mayank Kulkarni, Daniel Preotiuc-PietroACL 2020 · 31 citations
- Cross-lingual Spoken Language Understanding with Regularized Representation AlignmentZihan Liu, Genta Indra Winata, Peng Xu, Zhaojiang Lin et al.EMNLP 2020 · 16 citations
Related papers
- Three Heads Are Better than One: Improving Cross-Domain NER with Progressive Decomposed NetworkXuming Hu, Zhaochen Hong, Yong Jiang, Zhichao Lin et al.AAAI 2024 · 1 citation
- PDALN: Progressive Domain Adaptation over a Pre-trained Model for Low-Resource Cross-Domain Named Entity RecognitionTao Zhang, Congying Xia, Philip S. Yu, Zhiwei Liu et al.EMNLP 2021 · 22 citations
- Entity Extraction in Low Resource Domains with Selective Pre-training of Large Language ModelsAniruddha Mahapatra, Sharmila Reddy Nangi, Aparna Garimella, Anandhavelu NatarajanEMNLP 2022 · 4 citations
- A Simple Yet Effective Subsequence-Enhanced Approach for Cross-Domain NERJinpeng Hu, Dandan Guo, Yang Liu, Zhuo Li et al.AAAI 2023 · 11 citations
- DynamicNER: A Dynamic, Multilingual, and Fine-Grained Dataset for LLM-based Named Entity RecognitionHanjun Luo, Yingbin Jin, Yiran Wang, Xinfeng Li et al.EMNLP 2025
