Multilingual Transfer Learning for QA using Translation as Data Augmentation
Mihaela A. Bornea, Lin Pan, Sara Rosenthal, Radu Florian, Avirup Sil
Abstract
Prior work on multilingual question answering has mostly focused on using large multilingual pre-trained language models (LM) to perform zero-shot language-wise learning: train a QA model on English and test on other languages. In this work, we explore strategies that improve cross-lingual transfer by bringing the multilingual embeddings closer in the semantic space. Our first strategy augments the original English training data with machine translation-generated data. This results in a corpus of multilingual silver-labeled QA pairs that is 14 times larger than the original training set. In addition, we propose two novel strategies, language adversarial training and language arbitration framework, which significantly improve the (zero-resource) cross-lingual transfer performance and result in LM embeddings that are less language-variant. Empirically, we show that the proposed models outperform the previous zero-shot baseline on the recently introduced multilingual MLQA and TyDiQA datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 05aadd99-2d7c-465b-8c58-6fe02429f662Cited by top-tier papers7
- Enhancing Cross-lingual Transfer by Manifold MixupHuiyun Yang, Huadong Chen, Hao Zhou, Lei LiICLR 2022 · 49 citations
- Constrained Decoding for Cross-lingual Label ProjectionDuong Minh Le, Yang Chen, Alan Ritter, Wei XuICLR 2024 · 14 citations
- Multi-Aspect Heterogeneous Graph AugmentationYuchen Zhou, Yanan Cao, Yongchao Liu, Yanmin Shang et al.WWW 2023 · 6 citations
- Understanding LLMs' Cross-Lingual Context Retrieval: How Good It Is And Where It Comes FromChangjiang Gao, Hankun Lin, Xin Huang, Xue Han et al.EMNLP 2025 · 1 citation
- Aya Model: An Instruction Finetuned Open-Access Multilingual Language ModelAhmet Üstün, Viraat Aryabumi, Zheng Xin Yong, Wei-Yin Ko et al.ACL 2024
Builds on5
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- FreeLB: Enhanced Adversarial Training for Natural Language UnderstandingChen Zhu, Yu Cheng, Zhe Gan, Siqi Sun et al.ICLR 2020 · 502 citations
- On the Cross-lingual Transferability of Monolingual RepresentationsMikel Artetxe, Sebastian Ruder, Dani YogatamaACL 2020 · 57 citations
- MLQA: Evaluating Cross-lingual Extractive Question AnsweringPatrick Lewis, Barlas Oguz, Ruty Rinott, Sebastian Riedel et al.ACL 2020 · 52 citations
- Enhancing Answer Boundary Detection for Multilingual Machine Reading ComprehensionFei Yuan, Linjun Shou, Xuanyu Bai, Ming Gong et al.ACL 2020 · 21 citations
Related papers
- Improving the Cross-Lingual Generalisation in Visual Question AnsweringFarhad Nooralahzadeh, Rico SennrichAAAI 2023 · 8 citations
- Learning Disentangled Semantic Representations for Zero-Shot Cross-Lingual Transfer in Multilingual Machine Reading ComprehensionLinjuan Wu, Shaojuan Wu, Xiaowang Zhang, Deyi Xiong et al.ACL 2022 · 18 citations
- Improving Zero-Shot Cross-Lingual Transfer Learning via Robust TrainingKuan-Hao Huang, Wasi Uddin Ahmad, Nanyun Peng, Kai-Wei ChangEMNLP 2021 · 29 citations
- LAReQA: Language-Agnostic Answer Retrieval from a Multilingual PoolUma Roy, Noah Constant, Rami Al-Rfou, Aditya Barua et al.EMNLP 2020 · 39 citations
- From Zero to Hero: On the Limitations of Zero-Shot Language Transfer with Multilingual TransformersAnne Lauscher, Vinit Ravishankar, Ivan Vulic, Goran GlavasEMNLP 2020 · 235 citations
