The Secret is in the Spectra: Predicting Cross-lingual Task Performance with Spectral Similarity Measures
Haim Dubossarsky, Ivan Vulic, Roi Reichart, Anna Korhonen
Abstract
Performance in cross-lingual NLP tasks is impacted by the (dis)similarity of languages at hand: e.g., previous work has suggested there is a connection between the expected success of bilingual lexicon induction (BLI) and the assumption of (approximate) isomorphism between monolingual embedding spaces. In this work we present a large-scale study focused on the correlations between monolingual embedding space similarity and task performance, covering thousands of language pairs and four different tasks: BLI, parsing, POS tagging and MT. We hypothesize that statistics of the spectrum of each monolingual embedding space indicate how well they can be aligned. We then introduce several isomorphism measures between two embedding spaces, based on the relevant statistics of their individual spectra. We empirically show that 1) language similarity scores derived from such spectral isomorphism measures are strongly associated with performance observed in different crosslingual tasks, and 2) our spectral-based measures consistently outperform previous standard isomorphism measures, while being computationally more tractable and easier to interpret. Finally, our measures capture complementary information to typologically driven language distance measures, and the combination of measures from the two families yields even higher task performance correlations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Spectral Editing of Activations for Large Language Model AlignmentYifu Qiu, Zheng Zhao, Yftah Ziser, Anna Korhonen et al.NeurIPS 2024 · 66 citations
- Improving Word Translation via Two-Stage Contrastive LearningYaoyiran Li, Fangyu Liu, Nigel Collier, Anna Korhonen et al.ACL 2022 · 32 citations
- Are All Good Word Vector Spaces Isomorphic?Ivan Vulic, Sebastian Ruder, Anders SøgaardEMNLP 2020 · 7 citations
- IsoVec: Controlling the Relative Isomorphism of Word Embedding SpacesKelly Marchisio, Neha Verma, Kevin Duh, Philipp KoehnEMNLP 2022 · 6 citations
- A Massively Multilingual Analysis of Cross-linguality in Shared Embedding SpaceAlexander Jones, William Yang Wang, Kyle MahowaldEMNLP 2021 · 4 citations
Builds on2
Related papers
- Revisiting the Context Window for Cross-lingual Word EmbeddingsRyokan Ri, Yoshimasa TsuruokaACL 2020 · 4 citations
- LNMap: Departures from Isomorphic Assumption in Bilingual Lexicon Induction Through Non-Linear Mapping in Latent SpaceTasnim Mohiuddin, M. Saiful Bari, Shafiq Rayhan JotyEMNLP 2020 · 13 citations
- Filtered Inner Product Projection for Crosslingual Embedding AlignmentVin Sachidananda, Ziyi Yang, Chenguang ZhuICLR 2021 · 13 citations
- Bridging Linguistic Typology and Multilingual Machine Translation with Multi-View Language RepresentationsArturo Oncevay, Barry Haddow, Alexandra BirchEMNLP 2020 · 1 citation
- Enhancing Bilingual Lexicon Induction via Bi-directional Translation Pair RetrievingQiuyu Ding, Hailong Cao, Tiejun ZhaoAAAI 2024 · 3 citations
