Wider & Closer: Mixture of Short-channel Distillers for Zero-shot Cross-lingual Named Entity Recognition
Jun-Yu Ma, Beiduo Chen, Jia-Chen Gu, Zhenhua Ling, Wu Guo, Quan Liu, Zhigang Chen, Cong Liu
Abstract
Zero-shot cross-lingual named entity recognition (NER) aims at transferring knowledge from annotated and rich-resource data in source languages to unlabeled and lean-resource data in target languages. Existing mainstream methods based on the teacher-student distillation framework ignore the rich and complementary information lying in the intermediate layers of pre-trained language models, and domain-invariant information is easily lost during transfer. In this study, a mixture of short-channel distillers (MSD) method is proposed to fully interact the rich hierarchical information in the teacher model and to transfer knowledge to the student model sufficiently and efficiently. Concretely, a multi-channel distillation framework is designed for sufficient information transfer by aggregating multiple distillers as a mixture. Besides, an unsupervised method adopting parallel domain adaptation is proposed to shorten the channels between the teacher and student models to preserve domain-invariant features. Experiments on four datasets across nine languages demonstrate that the proposed method achieves new state-of-the-art performance on zero-shot cross-lingual NER and shows great generalization and compatibility across languages and fields.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ebfcd1f3-e266-49c6-87ec-d7bcfe2ea213Cited by top-tier papers4
- Discrepancy and Uncertainty Aware Denoising Knowledge Distillation for Zero-Shot Cross-Lingual Named Entity RecognitionLing Ge, Chunming Hu, Guanghui Ma, Jihong Liu et al.AAAI 2024 · 9 citations
- Neighboring Perturbations of Knowledge Editing on Large Language ModelsJun-Yu Ma, Zhen-Hua Ling, Ningyu Zhang, Jia-Chen GuICML 2024 · 6 citations
- DA-Net: A Disentangled and Adaptive Network for Multi-Source Cross-Lingual Transfer LearningLing Ge, Chunming Hu, Guanghui Ma, Jihong Liu et al.AAAI 2024 · 3 citations
- Large Margin Representation Learning for Robust Cross-lingual Named Entity RecognitionGuangcheng Zhu, Ruixuan Xiao, Haobo Wang, Zhen Zhu et al.ACL 2025 · 1 citation
Builds on7
- SMART: Robust and Efficient Fine-Tuning for Pre-trained Natural Language Models through Principled Regularized OptimizationHaoming Jiang, Pengcheng He, Weizhu Chen, Xiaodong Liu et al.ACL 2020 · 148 citations
- Entity Enhanced BERT Pre-training for Chinese NERChen Jia, Yuefeng Shi, Qinrong Yang, Yue ZhangEMNLP 2020 · 59 citations
- Single-/Multi-Source Cross-Lingual NER via Teacher-Student Learning on Unlabeled Data in Target LanguageQianhui Wu, Zijia Lin, Börje Karlsson, Jianguang Lou et al.ACL 2020 · 59 citations
- Improving Zero-Shot Cross-Lingual Transfer Learning via Robust TrainingKuan-Hao Huang, Wasi Uddin Ahmad, Nanyun Peng, Kai-Wei ChangEMNLP 2021 · 29 citations
- An Unsupervised Multiple-Task and Multiple-Teacher Model for Cross-lingual Named Entity RecognitionZhuoran Li, Chunming Hu, Xiaohui Guo, Junfan Chen et al.ACL 2022 · 23 citations
Related papers
- ProKD: An Unsupervised Prototypical Knowledge Distillation Network for Zero-Resource Cross-Lingual Named Entity RecognitionLing Ge, Chunming Hu, Guanghui Ma, Hong Zhang et al.AAAI 2023 · 9 citations
- PDALN: Progressive Domain Adaptation over a Pre-trained Model for Low-Resource Cross-Domain Named Entity RecognitionTao Zhang, Congying Xia, Philip S. Yu, Zhiwei Liu et al.EMNLP 2021 · 22 citations
- XtremeDistil: Multi-stage Distillation for Massive Multilingual ModelsSubhabrata Mukherjee, Ahmed Hassan AwadallahACL 2020 · 4 citations
- Pretrained Bidirectional Distillation for Machine TranslationYimeng Zhuang, Mei TuACL 2023 · 3 citations
- Knowledge Distillation for Large Language Models through Residual LearningThinh On, Hengzhi Pei, Leonard Lausen, George KarypisICLR 2026 · 5 citations
