Bridging Linguistic Typology and Multilingual Machine Translation with Multi-View Language Representations
Arturo Oncevay, Barry Haddow, Alexandra Birch
摘要
Sparse language vectors from linguistic typology databases and learned embeddings from tasks like multilingual machine translation have been investigated in isolation, without analysing how they could benefit from each other's language characterisation. We propose to fuse both views using singular vector canonical correlation analysis and study what kind of information is induced from each source. By inferring typological features and language phylogenies, we observe that our representations embed typology and strengthen correlations with language relationships. We then take advantage of our multi-view language vector space for multilingual machine translation, where we achieve competitive overall translation accuracy in tasks that require information about language similarities, such as language clustering and ranking candidates for multilingual transfer. With our method, which is also released as a tool, we can easily project and assess new languages without expensive retraining of massive multilingual or ranking models, which are major disadvantages of related approaches.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Adaptive Token-level Cross-lingual Feature Mixing for Multilingual Neural Machine TranslationJunpeng Liu, Kaiyu Huang, Jiuyi Li, Huan Liu 等EMNLP 2022 · 被引用 5 次
- MMTE: Corpus and Metrics for Evaluating Machine Translation Quality of Metaphorical LanguageShun Wang, Ge Zhang, Han Wu, Tyler Loakman 等EMNLP 2024 · 被引用 3 次
- Beyond Demographics: Enhancing Cultural Value Survey Simulation with Multi-Stage Personality-Driven Cognitive ReasoningHaijiang Liu, Qiyuan Li, Chao Gao, Yong Cao 等EMNLP 2025
- GradSim: Gradient-Based Language Grouping for Effective Multilingual TrainingMingyang Wang, Heike Adel, Lukas Lange, Jannik Strötgen 等EMNLP 2023
- Can LLMs Really Learn to Translate a Low-Resource Language from One Grammar Book?Seth Aycock, David Stap, Di Wu, Christof Monz 等ICLR 2025
它引用的顶会 Paper1
相关 Paper
- The Secret is in the Spectra: Predicting Cross-lingual Task Performance with Spectral Similarity MeasuresHaim Dubossarsky, Ivan Vulic, Roi Reichart, Anna KorhonenEMNLP 2020
- Analysis of Multi-Source Language Training in Cross-Lingual TransferSeong Hoon Lim, Taejun Yun, Jinhyeon Kim, Jihun Choi 等ACL 2024 · 被引用 1 次
- Normalization of Language Embeddings for Cross-Lingual AlignmentPrince Osei Aboagye, Yan Zheng, Chin-Chia Michael Yeh, Junpeng Wang 等ICLR 2022 · 被引用 12 次
- Discovering Low-rank Subspaces for Language-agnostic Multilingual RepresentationsZhihui Xie, Handong Zhao, Tong Yu, Shuai LiEMNLP 2022 · 被引用 3 次
- T-FREE: Subword Tokenizer-Free Generative LLMs via Sparse Representations for Memory-Efficient EmbeddingsBjörn Deiseroth, Manuel Brack, Patrick Schramowski, Kristian Kersting 等EMNLP 2024 · 被引用 2 次
