Synthetic Data Augmentation for Zero-Shot Cross-Lingual Question Answering
Arij Riabi, Thomas Scialom, Rachel Keraron, Benoît Sagot, Djamé Seddah, Jacopo Staiano
Abstract
Coupled with the availability of large scale datasets, deep learning architectures have enabled rapid progress on Question Answering tasks. However, most of those datasets are in English, and the performances of state-of-theart multilingual models are significantly lower when evaluated on non-English data. Due to high data collection costs, it is not realistic to obtain annotated data for each language one desires to support. We propose a method to improve Crosslingual Question Answering performance without requiring additional annotated data, leveraging Question Generation models to produce synthetic samples in a cross-lingual fashion. We show that the proposed method allows to significantly outperform the baselines trained on English data only, establishing thus a new state-of-the-art on four multilingual datasets: MLQA, XQuAD, SQuAD-it and PIAF (fr). * * : equal contribution. The work of Arij Riabi was partly carried out while she was working at reciTAL. 1 https://rajpurkar.github.io/ SQuAD-explorer/ et al. (2020) and Lewis et al. ( 2020a ) concurrently proposed two different evaluation sets which are comparable to the SQuAD development set. Both reach the same conclusion: due to the lack of non-English training data, models do not achieve the same performance in Non-English languages than they do in English. To the best of our knowledge, no method has been proposed to fill this gap.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- One Question Answering Model for Many Languages with Cross-lingual Dense Passage RetrievalAkari Asai, Xinyan Yu, Jungo Kasai, Hanna HajishirziNeurIPS 2021 · 86 citations
- Fine-tuned Language Models are Continual LearnersThomas Scialom, Tuhin Chakrabarty, Smaranda MuresanEMNLP 2022 · 46 citations
- Building a Foundational Guardrail for General Agentic Systems via Synthetic DataYue Huang, Hang Hua, Yujun Zhou, Pengcheng Jing et al.ICLR 2026 · 29 citations
- IDK-MRC: Unanswerable Questions for Indonesian Machine Reading ComprehensionRifki Afina Putri, Alice OhEMNLP 2022 · 11 citations
- ChemOrch: Empowering LLMs with Chemical Intelligence via Groundbreaking Synthetic InstructionsYue Huang, Zhengzhe Jiang, Xiaonan Luo, Kehan Guo et al.NeurIPS 2025 · 5 citations
Builds on6
- MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained TransformersWenhui Wang, Furu Wei, Li Dong, Hangbo Bao et al.NeurIPS 2020 · 2,727 citations
- CamemBERT: a Tasty French Language ModelLouis Martin, Benjamin Muller, Pedro Javier Ortiz Suárez, Yoann Dupont et al.ACL 2020 · 703 citations
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- Cross-Lingual Natural Language Generation via Pre-TrainingZewen Chi, Li Dong, Furu Wei, Wenhui Wang et al.AAAI 2020 · 142 citations
- On the Cross-lingual Transferability of Monolingual RepresentationsMikel Artetxe, Sebastian Ruder, Dani YogatamaACL 2020 · 57 citations
Related papers
- Cross-lingual Transfer for Automatic Question Generation by Learning Interrogative Structures in Target LanguagesSeonjeong Hwang, Yunsu Kim, Gary Geunbae LeeEMNLP 2024 · 2 citations
- Generative Language Models for Paragraph-Level Question GenerationAsahi Ushio, Fernando Alva-Manchego, José Camacho-ColladosEMNLP 2022 · 30 citations
- Multilingual Transfer Learning for QA using Translation as Data AugmentationMihaela A. Bornea, Lin Pan, Sara Rosenthal, Radu Florian et al.AAAI 2021 · 45 citations
- MLQA: Evaluating Cross-lingual Extractive Question AnsweringPatrick Lewis, Barlas Oguz, Ruty Rinott, Sebastian Riedel et al.ACL 2020 · 52 citations
- M4-RAG: A Massive-Scale Multilingual Multi-Cultural Multimodal RAGDavid Anugraha, Patrick Amadeus Irawan, Anshul Singh, En-Shiun Annie Lee et al.CVPR 2026 · 2 citations
