Bridging Linguistic Typology and Multilingual Machine Translation with Multi-View Language Representations
Arturo Oncevay, Barry Haddow, Alexandra Birch
Abstract
Sparse language vectors from linguistic typology databases and learned embeddings from tasks like multilingual machine translation have been investigated in isolation, without analysing how they could benefit from each other's language characterisation. We propose to fuse both views using singular vector canonical correlation analysis and study what kind of information is induced from each source. By inferring typological features and language phylogenies, we observe that our representations embed typology and strengthen correlations with language relationships. We then take advantage of our multi-view language vector space for multilingual machine translation, where we achieve competitive overall translation accuracy in tasks that require information about language similarities, such as language clustering and ranking candidates for multilingual transfer. With our method, which is also released as a tool, we can easily project and assess new languages without expensive retraining of massive multilingual or ranking models, which are major disadvantages of related approaches.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cac87f5a-33cc-4221-8798-736295303f01Cited by top-tier papers5
- Adaptive Token-level Cross-lingual Feature Mixing for Multilingual Neural Machine TranslationJunpeng Liu, Kaiyu Huang, Jiuyi Li, Huan Liu et al.EMNLP 2022 · 5 citations
- MMTE: Corpus and Metrics for Evaluating Machine Translation Quality of Metaphorical LanguageShun Wang, Ge Zhang, Han Wu, Tyler Loakman et al.EMNLP 2024 · 3 citations
- Beyond Demographics: Enhancing Cultural Value Survey Simulation with Multi-Stage Personality-Driven Cognitive ReasoningHaijiang Liu, Qiyuan Li, Chao Gao, Yong Cao et al.EMNLP 2025
- GradSim: Gradient-Based Language Grouping for Effective Multilingual TrainingMingyang Wang, Heike Adel, Lukas Lange, Jannik Strötgen et al.EMNLP 2023
- Can LLMs Really Learn to Translate a Low-Resource Language from One Grammar Book?Seth Aycock, David Stap, Di Wu, Christof Monz et al.ICLR 2025
Builds on1
Related papers
- The Secret is in the Spectra: Predicting Cross-lingual Task Performance with Spectral Similarity MeasuresHaim Dubossarsky, Ivan Vulic, Roi Reichart, Anna KorhonenEMNLP 2020
- Analysis of Multi-Source Language Training in Cross-Lingual TransferSeong Hoon Lim, Taejun Yun, Jinhyeon Kim, Jihun Choi et al.ACL 2024 · 1 citation
- Normalization of Language Embeddings for Cross-Lingual AlignmentPrince Osei Aboagye, Yan Zheng, Chin-Chia Michael Yeh, Junpeng Wang et al.ICLR 2022 · 12 citations
- Discovering Low-rank Subspaces for Language-agnostic Multilingual RepresentationsZhihui Xie, Handong Zhao, Tong Yu, Shuai LiEMNLP 2022 · 3 citations
- T-FREE: Subword Tokenizer-Free Generative LLMs via Sparse Representations for Memory-Efficient EmbeddingsBjörn Deiseroth, Manuel Brack, Patrick Schramowski, Kristian Kersting et al.EMNLP 2024 · 2 citations
