Implicit Word Reordering with Knowledge Distillation for Cross-Lingual Dependency Parsing
Zhuoran Li, Chunming Hu, Junfan Chen, Zhijun Chen, Richong Zhang
摘要
Word order difference between source and target languages is a major obstacle to cross-lingual transfer, especially in the dependency parsing task. Current works are mostly based on order-agnostic models or word reordering to mitigate this problem. However, such methods either do not leverage grammatical information naturally contained in word order or are computationally expensive as the permutation space grows exponentially with the sentence length. Moreover, the reordered source sentence with an unnatural word order may be a form of noising that harms the model learning. To this end, we propose an Implicit Word Reordering framework with Knowledge Distillation (IWR-KD). This framework is inspired by that deep networks are good at learning feature linearization corresponding to meaningful data transformation, e.g. word reordering. To realize this idea, we introduce a knowledge distillation framework composed of a word-reordering teacher model and a dependency parsing student model. We verify our proposed method on Universal Dependency Treebanks across 31 different languages and show it outperforms a series of competitors, together with experimental analysis to illustrate how our method works towards training a robust parser.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- On the Importance of Word Order Information in Cross-lingual Sequence LabelingZihan Liu, Genta Indra Winata, Samuel Cahyawijaya, Andrea Madotto 等AAAI 2021 · 被引用 29 次
- Substructure Distribution Projection for Zero-Shot Cross-Lingual Dependency ParsingFreda Shi, Kevin Gimpel, Karen LivescuACL 2022 · 被引用 10 次
- Neural Syntactic Preordering for Controlled Paraphrase GenerationTanya Goyal, Greg DurrettACL 2020 · 被引用 9 次
- Word Reordering for Zero-shot Cross-lingual Structured PredictionTao Ji, Yong Jiang, Tao Wang, Zhongqiang Huang 等EMNLP 2021 · 被引用 3 次
相关 Paper
- ProKD: An Unsupervised Prototypical Knowledge Distillation Network for Zero-Resource Cross-Lingual Named Entity RecognitionLing Ge, Chunming Hu, Guanghui Ma, Hong Zhang 等AAAI 2023 · 被引用 9 次
- Domain-Adapted Dependency Parsing for Cross-Domain Named Entity RecognitionChenxiao Dou, Xianghui Sun, Yaoshu Wang, Yunjie Ji 等AAAI 2023 · 被引用 8 次
- Fine-Grained Analysis of Cross-Linguistic Syntactic DivergencesDmitry Nikolaev, Ofir Arviv, Taelin Karidi, Neta Kenneth 等ACL 2020 · 被引用 1 次
- Discrepancy and Uncertainty Aware Denoising Knowledge Distillation for Zero-Shot Cross-Lingual Named Entity RecognitionLing Ge, Chunming Hu, Guanghui Ma, Jihong Liu 等AAAI 2024 · 被引用 9 次
- Towards Understanding and Improving Knowledge Distillation for Neural Machine TranslationSongming Zhang, Yunlong Liang, Shuaibo Wang, Yufeng Chen 等ACL 2023 · 被引用 8 次
