Discrepancy and Uncertainty Aware Denoising Knowledge Distillation for Zero-Shot Cross-Lingual Named Entity Recognition
Ling Ge, Chunming Hu, Guanghui Ma, Jihong Liu, Hong Zhang
摘要
The knowledge distillation-based approaches have recently yielded state-of-the-art (SOTA) results for cross-lingual NER tasks in zero-shot scenarios. These approaches typically employ a teacher network trained with the labelled source (rich-resource) language to infer pseudo-soft labels for the unlabelled target (zero-shot) language, and force a student network to approximate these pseudo labels to achieve knowledge transfer. However, previous works have rarely discussed the issue of pseudo-label noise caused by the source-target language gap, which can mislead the training of the student network and result in negative knowledge transfer. This paper proposes an discrepancy and uncertainty aware Denoising Knowledge Distillation model (DenKD) to tackle this issue. Specifically, DenKD uses a discrepancy-aware denoising representation learning method to optimize the class representations of the target language produced by the teacher network, thus enhancing the quality of pseudo labels and reducing noisy predictions. Further, DenKD employs an uncertainty-aware denoising method to quantify the pseudo-label noise and adjust the focus of the student network on different samples during knowledge distillation, thereby mitigating the noise's adverse effects. We conduct extensive experiments on 28 languages including 4 languages not covered by the pre-trained models, and the results demonstrate the effectiveness of our DenKD.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper14
- Your Classifier can Secretly Suffice Multi-Source Domain AdaptationNaveen Venkat, Jogendra Nath Kundu, Durgesh Kumar Singh, Ambareesh Revanur 等NeurIPS 2020 · 被引用 95 次
- Enhanced Meta-Learning for Cross-Lingual Named Entity Recognition with Minimal ResourcesQianhui Wu, Zijia Lin, Guoxin Wang, Hui Chen 等AAAI 2020 · 被引用 72 次
- Multi-Granularity Structural Knowledge Distillation for Language Model CompressionChang Liu, Chongyang Tao, Jiazhan Feng, Dongyan ZhaoACL 2022 · 被引用 64 次
- Single-/Multi-Source Cross-Lingual NER via Teacher-Student Learning on Unlabeled Data in Target LanguageQianhui Wu, Zijia Lin, Börje Karlsson, Jianguang Lou 等ACL 2020 · 被引用 59 次
- Zero-Resource Cross-Lingual Named Entity RecognitionM. Saiful Bari, Shafiq R. Joty, Prathyusha JwalapuramAAAI 2020 · 被引用 55 次
相关 Paper
- ProKD: An Unsupervised Prototypical Knowledge Distillation Network for Zero-Resource Cross-Lingual Named Entity RecognitionLing Ge, Chunming Hu, Guanghui Ma, Hong Zhang 等AAAI 2023 · 被引用 9 次
- Wider & Closer: Mixture of Short-channel Distillers for Zero-shot Cross-lingual Named Entity RecognitionJun-Yu Ma, Beiduo Chen, Jia-Chen Gu, Zhenhua Ling 等EMNLP 2022 · 被引用 3 次
- ConNER: Consistency Training for Cross-lingual Named Entity RecognitionRan Zhou, Xin Li, Lidong Bing, Erik Cambria 等EMNLP 2022 · 被引用 16 次
- CoLaDa: A Collaborative Label Denoising Framework for Cross-lingual Named Entity RecognitionTingting Ma, Qianhui Wu, Huiqiang Jiang, Börje Karlsson 等ACL 2023 · 被引用 5 次
- Knowledge Diffusion for DistillationTao Huang, Yuan Zhang, Mingkai Zheng, Shan You 等NeurIPS 2023 · 被引用 125 次
