Cross-domain NER with Generated Task-Oriented Knowledge: An Empirical Study from Information Density Perspective
Zhihao Zhang, Sophia Yat Mei Lee, Junshuang Wu, Dong Zhang, Shoushan Li, Erik Cambria, Guodong Zhou
摘要
Cross-domain Named Entity Recognition (CD-NER) is crucial for Knowledge Graph (KG) construction and natural language processing (NLP), enabling learning from source to target domains with limited data. Previous studies often rely on manually collected entity-relevant sentences from the web or attempt to bridge the gap between tokens and entity labels across domains. These approaches are time-consuming and inefficient, as these data are often weakly correlated with the target task and require extensive pre-training. To address these issues, we propose automatically generating task-oriented knowledge (GTOK) using large language models (LLMs), focusing on the reasoning process of entity extraction. Then, we employ taskoriented pre-training (TOPT) to facilitate domain adaptation. Additionally, current crossdomain NER methods often lack explicit explanations for their effectiveness. Therefore, we introduce the concept of information density to better evaluate the model's effectiveness before performing entity recognition. We conduct systematic experiments and analyses to demonstrate the effectiveness of our proposed approach and the validity of using information density for model evaluation † .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 被引用 5,863 次
- CrossNER: Evaluating Cross-Domain Named Entity RecognitionZihan Liu, Yan Xu, Tiezheng Yu, Wenliang Dai 等AAAI 2021 · 被引用 201 次
- Few-shot Slot Tagging with Collapsed Dependency Transfer and Label-enhanced Task-adaptive Projection NetworkYutai Hou, Wanxiang Che, Yongkui Lai, Zhihan Zhou 等ACL 2020 · 被引用 186 次
- Joint Multi-modal Aspect-Sentiment Analysis with Auxiliary Cross-modal Relation DetectionXincheng Ju, Dong Zhang, Rong Xiao, Junhui Li 等EMNLP 2021 · 被引用 130 次
- Synthetic Data Generation with Large Language Models for Text Classification: Potential and LimitationsZhuoyan Li, Hangxiao Zhu, Zhuoran Lu, Ming YinEMNLP 2023 · 被引用 102 次
相关 Paper
- Domain-oriented Language Modeling with Adaptive Hybrid Masking and Optimal Transport AlignmentDenghui Zhang, Zixuan Yuan, Yanchi Liu, Hao Liu 等KDD 2021 · 被引用 7 次
- Three Heads Are Better than One: Improving Cross-Domain NER with Progressive Decomposed NetworkXuming Hu, Zhaochen Hong, Yong Jiang, Zhichao Lin 等AAAI 2024 · 被引用 1 次
- AdvPicker: Effectively Leveraging Unlabeled Data via Adversarial Discriminator for Cross-Lingual NERWeile Chen, Huiqiang Jiang, Qianhui Wu, Börje Karlsson 等ACL 2021
- Exploiting Structured Knowledge in Text via Graph-Guided Representation LearningTao Shen, Yi Mao, Pengcheng He, Guodong Long 等EMNLP 2020 · 被引用 60 次
- Entity Extraction in Low Resource Domains with Selective Pre-training of Large Language ModelsAniruddha Mahapatra, Sharmila Reddy Nangi, Aparna Garimella, Anandhavelu NatarajanEMNLP 2022 · 被引用 4 次
