Should All Cross-Lingual Embeddings Speak English?
Antonios Anastasopoulos, Graham Neubig
Abstract
Most of recent work in cross-lingual word embeddings is severely Anglocentric. The vast majority of lexicon induction evaluation dictionaries are between English and another language, and the English embedding space is selected by default as the hub when learning in a multilingual setting. With this work, however, we challenge these practices. First, we show that the choice of hub language can significantly impact downstream lexicon induction and zero-shot POS tagging performance. Second, we both expand a standard Englishcentered evaluation dictionary collection to include all language pairs using triangulation, and create new dictionaries for under-represented languages. 1 Evaluating established methods over all these language pairs sheds light into their suitability for aligning embeddings from distant languages and presents new challenges for the field. Finally, in our analysis we identify general guidelines for strong cross-lingual embedding baselines, that extend to language pairs that do not include English.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e571c8a2-e192-4619-8ac2-67f6351ab727Cited by top-tier papers3
- Predicting Performance for Natural Language Processing TasksMengzhou Xia, Antonios Anastasopoulos, Ruochen Xu, Yiming Yang et al.ACL 2020 · 39 citations
- XTREME-R: Towards More Challenging and Nuanced Multilingual EvaluationSebastian Ruder, Noah Constant, Jan A. Botha, Aditya Siddhant et al.EMNLP 2021 · 10 citations
- XL-WiC: A Multilingual Benchmark for Evaluating Semantic ContextualizationAlessandro Raganato, Tommaso Pasini, José Camacho-Collados, Mohammad Taher PilehvarEMNLP 2020 · 2 citations
Related papers
- Revisiting the Context Window for Cross-lingual Word EmbeddingsRyokan Ri, Yoshimasa TsuruokaACL 2020 · 4 citations
- Beyond Offline Mapping: Learning Cross-lingual Word Embeddings through Context AnchoringAitor Ormazabal, Mikel Artetxe, Aitor Soroa, Gorka Labaka et al.ACL 2021
- Enhancing Bilingual Lexicon Induction via Bi-directional Translation Pair RetrievingQiuyu Ding, Hailong Cao, Tiejun ZhaoAAAI 2024 · 3 citations
- IsoVec: Controlling the Relative Isomorphism of Word Embedding SpacesKelly Marchisio, Neha Verma, Kevin Duh, Philipp KoehnEMNLP 2022 · 6 citations
- A Call for More Rigor in Unsupervised Cross-lingual LearningMikel Artetxe, Sebastian Ruder, Dani Yogatama, Gorka Labaka et al.ACL 2020 · 5 citations
