Wider & Closer: Mixture of Short-channel Distillers for Zero-shot Cross-lingual Named Entity Recognition
Jun-Yu Ma, Beiduo Chen, Jia-Chen Gu, Zhenhua Ling, Wu Guo, Quan Liu, Zhigang Chen, Cong Liu
摘要
Zero-shot cross-lingual named entity recognition (NER) aims at transferring knowledge from annotated and rich-resource data in source languages to unlabeled and lean-resource data in target languages. Existing mainstream methods based on the teacher-student distillation framework ignore the rich and complementary information lying in the intermediate layers of pre-trained language models, and domain-invariant information is easily lost during transfer. In this study, a mixture of short-channel distillers (MSD) method is proposed to fully interact the rich hierarchical information in the teacher model and to transfer knowledge to the student model sufficiently and efficiently. Concretely, a multi-channel distillation framework is designed for sufficient information transfer by aggregating multiple distillers as a mixture. Besides, an unsupervised method adopting parallel domain adaptation is proposed to shorten the channels between the teacher and student models to preserve domain-invariant features. Experiments on four datasets across nine languages demonstrate that the proposed method achieves new state-of-the-art performance on zero-shot cross-lingual NER and shows great generalization and compatibility across languages and fields.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Discrepancy and Uncertainty Aware Denoising Knowledge Distillation for Zero-Shot Cross-Lingual Named Entity RecognitionLing Ge, Chunming Hu, Guanghui Ma, Jihong Liu 等AAAI 2024 · 被引用 9 次
- Neighboring Perturbations of Knowledge Editing on Large Language ModelsJun-Yu Ma, Zhen-Hua Ling, Ningyu Zhang, Jia-Chen GuICML 2024 · 被引用 6 次
- DA-Net: A Disentangled and Adaptive Network for Multi-Source Cross-Lingual Transfer LearningLing Ge, Chunming Hu, Guanghui Ma, Jihong Liu 等AAAI 2024 · 被引用 3 次
- Large Margin Representation Learning for Robust Cross-lingual Named Entity RecognitionGuangcheng Zhu, Ruixuan Xiao, Haobo Wang, Zhen Zhu 等ACL 2025 · 被引用 1 次
它引用的顶会 Paper7
- SMART: Robust and Efficient Fine-Tuning for Pre-trained Natural Language Models through Principled Regularized OptimizationHaoming Jiang, Pengcheng He, Weizhu Chen, Xiaodong Liu 等ACL 2020 · 被引用 148 次
- Entity Enhanced BERT Pre-training for Chinese NERChen Jia, Yuefeng Shi, Qinrong Yang, Yue ZhangEMNLP 2020 · 被引用 59 次
- Single-/Multi-Source Cross-Lingual NER via Teacher-Student Learning on Unlabeled Data in Target LanguageQianhui Wu, Zijia Lin, Börje Karlsson, Jianguang Lou 等ACL 2020 · 被引用 59 次
- Improving Zero-Shot Cross-Lingual Transfer Learning via Robust TrainingKuan-Hao Huang, Wasi Uddin Ahmad, Nanyun Peng, Kai-Wei ChangEMNLP 2021 · 被引用 29 次
- An Unsupervised Multiple-Task and Multiple-Teacher Model for Cross-lingual Named Entity RecognitionZhuoran Li, Chunming Hu, Xiaohui Guo, Junfan Chen 等ACL 2022 · 被引用 23 次
相关 Paper
- ProKD: An Unsupervised Prototypical Knowledge Distillation Network for Zero-Resource Cross-Lingual Named Entity RecognitionLing Ge, Chunming Hu, Guanghui Ma, Hong Zhang 等AAAI 2023 · 被引用 9 次
- PDALN: Progressive Domain Adaptation over a Pre-trained Model for Low-Resource Cross-Domain Named Entity RecognitionTao Zhang, Congying Xia, Philip S. Yu, Zhiwei Liu 等EMNLP 2021 · 被引用 22 次
- XtremeDistil: Multi-stage Distillation for Massive Multilingual ModelsSubhabrata Mukherjee, Ahmed Hassan AwadallahACL 2020 · 被引用 4 次
- Pretrained Bidirectional Distillation for Machine TranslationYimeng Zhuang, Mei TuACL 2023 · 被引用 3 次
- Knowledge Distillation for Large Language Models through Residual LearningThinh On, Hengzhi Pei, Leonard Lausen, George KarypisICLR 2026 · 被引用 5 次
