Model-Based Ranking of Source Languages for Zero-Shot Cross-Lingual Transfer
Abteen Ebrahimi, Adam Wiemerslage, Katharina von der Wense
Abstract
We present NN-RANK, an algorithm for ranking source languages for cross-lingual transfer, which leverages hidden representations from multilingual models and unlabeled targetlanguage data. We experiment with two pretrained multilingual models and two tasks: partof-speech tagging (POS) and named entity recognition (NER). We consider 51 source languages and evaluate on 56 and 72 target languages for POS and NER, respectively. When using in-domain data, NN-RANK beats stateof-the-art baselines that leverage lexical and linguistic features, with average improvements of up to 35.56 NDCG for POS and 18.14 NDCG for NER. As prior approaches can fall back to language-level features if target language data is not available, we show that NN-RANK remains competitive using only the Bible, an out-of-domain corpus available for a large number of languages. Ablations on the amount of unlabeled target data show that, for subsets consisting of as few as 25 examples, NN-RANK produces high-quality rankings which achieve 92.8% of the NDCG achieved using all available target data for ranking.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9c3c16eb-e503-4c8d-a032-e8c22d6836e6Builds on10
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- Cross-Lingual Ability of Multilingual BERT: An Empirical StudyKarthikeyan K, Zihan Wang, Stephen Mayhew, Dan RothICLR 2020 · 378 citations
- Emerging Cross-lingual Structure in Pretrained Language ModelsAlexis Conneau, Shijie Wu, Haoran Li, Luke Zettlemoyer et al.ACL 2020 · 210 citations
- Make the Best of Cross-lingual Transfer: Evidence from POS Tagging with over 100 LanguagesWietse de Vries, Martijn Wieling, Malvina NissimACL 2022 · 63 citations
- Single-/Multi-Source Cross-Lingual NER via Teacher-Student Learning on Unlabeled Data in Target LanguageQianhui Wu, Zijia Lin, Börje Karlsson, Jianguang Lou et al.ACL 2020 · 59 citations
Related papers
- Improving Low-Resource Languages in Pre-Trained Multilingual Language ModelsViktor Hangya, Hossain Shaikh Saadi, Alexander FraserEMNLP 2022 · 17 citations
- Unsupervised Cross-Lingual Part-of-Speech Tagging for Truly Low-Resource ScenariosRamy Eskander, Smaranda Muresan, Michael CollinsEMNLP 2020 · 16 citations
- Zero-Resource Cross-Lingual Named Entity RecognitionM. Saiful Bari, Shafiq R. Joty, Prathyusha JwalapuramAAAI 2020 · 55 citations
- How to Adapt Your Pretrained Multilingual Model to 1600 LanguagesAbteen Ebrahimi, Katharina KannACL 2021
- XLM-K: Improving Cross-Lingual Language Model Pre-training with Multilingual KnowledgeXiaoze Jiang, Yaobo Liang, Weizhu Chen, Nan DuanAAAI 2022 · 31 citations
