A Comparison of Architectures and Pretraining Methods for Contextualized Multilingual Word Embeddings
Niels van der Heijden, Samira Abnar, Ekaterina Shutova
摘要
The lack of annotated data in many languages is a well-known challenge within the field of multilingual natural language processing (NLP). Therefore, many recent studies focus on zero-shot transfer learning and joint training across languages to overcome data scarcity for low-resource languages. In this work we (i) perform a comprehensive comparison of state-of-the-art multilingual word and sentence encoders on the tasks of named entity recognition (NER) and part of speech (POS) tagging; and (ii) propose a new method for creating multilingual contextualized word embeddings, compare it to multiple baselines and show that it performs at or above state-of-the-art level in zero-shot transfer settings. Finally, we show that our method allows for better knowledge sharing across languages in a joint training setting.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
相关 Paper
- Zero-Resource Cross-Lingual Named Entity RecognitionM. Saiful Bari, Shafiq R. Joty, Prathyusha JwalapuramAAAI 2020 · 被引用 55 次
- Improving Low-Resource Languages in Pre-Trained Multilingual Language ModelsViktor Hangya, Hossain Shaikh Saadi, Alexander FraserEMNLP 2022 · 被引用 17 次
- Graph-Based Multilingual Label Propagation for Low-Resource Part-of-Speech TaggingAyyoob Imani, Silvia Severini, Masoud Jalili Sabet, François Yvon 等EMNLP 2022 · 被引用 8 次
- Multilingual Alignment of Contextual Word RepresentationsSteven Cao, Nikita Kitaev, Dan KleinICLR 2020 · 被引用 211 次
- On the Importance of Word Order Information in Cross-lingual Sequence LabelingZihan Liu, Genta Indra Winata, Samuel Cahyawijaya, Andrea Madotto 等AAAI 2021 · 被引用 29 次
