Overcoming Catastrophic Forgetting in Zero-Shot Cross-Lingual Generation
Tu Vu, Aditya Barua, Brian Lester, Daniel Cer, Mohit Iyyer, Noah Constant
Abstract
In this paper, we explore the challenging problem of performing a generative task in a target language when labeled data is only available in English, using summarization as a case study. We assume a strict setting with no access to parallel data or machine translation and find that common transfer learning approaches struggle in this setting, as a generative multilingual model fine-tuned purely on English catastrophically forgets how to generate non-English. Given the recent rise of parameter-efficient adaptation techniques, we conduct the first investigation into how one such method, prompt tuning (Lester et al., 2021), can overcome catastrophic forgetting to enable zero-shot cross-lingual generation. Our experiments show that parameter-efficient prompt tuning provides gains over standard fine-tuning when transferring between less-related languages, e.g., from English to Thai. However, a significant gap still remains between these methods and fully-supervised baselines. To improve cross-lingual transfer further, we explore several approaches, including: (1) mixing in unlabeled multilingual data, and (2) explicitly factoring prompts into recombinable language and task components. Our approaches can provide further quality gains, suggesting that robust zero-shot cross-lingual generation is within reach.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6e8943b1-9b0a-4594-b903-2c90e03358fdCited by top-tier papers14
- Pretraining Language Models with Human PreferencesTomasz Korbak, Kejian Shi, Angelica Chen, Rasika Vinayak Bhalerao et al.ICML 2023 · 287 citations
- Conditional Adapters: Parameter-efficient Transfer Learning with Fast InferenceTao Lei, Junwen Bai, Siddhartha Brahma, Joshua Ainslie et al.NeurIPS 2023 · 103 citations
- Middle-Layer Representation Alignment for Cross-Lingual Transfer in Fine-Tuned LLMsDanni Liu, Jan NiehuesACL 2025 · 23 citations
- Large Language Models Can Be Contextual Privacy Protection LearnersYijia Xiao, Yiqiao Jin, Yushi Bai, Yue Wu et al.EMNLP 2024 · 18 citations
- Multilingual Simplification of Medical TextsSebastian Joseph, Kathryn Kazanas, Keziah Reina, Vishnesh J. Ramanathan et al.EMNLP 2023 · 16 citations
Builds on19
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- Multitask Prompted Training Enables Zero-Shot Task GeneralizationVictor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach et al.ICLR 2022 · 1,976 citations
- Few-Shot Parameter-Efficient Fine-Tuning is Better and Cheaper than In-Context LearningHaokun Liu, Derek Tam, Mohammed Muqeeth, Jay Mohta et al.NeurIPS 2022 · 1,483 citations
- Towards a Unified View of Parameter-Efficient Transfer LearningJunxian He, Chunting Zhou, Xuezhe Ma, Taylor Berg-Kirkpatrick et al.ICLR 2022 · 1,182 citations
Related papers
- Efficient Unseen Language Adaptation for Multilingual Pre-Trained Language ModelsPo-Heng Chen, Yun-Nung ChenEMNLP 2024
- ADPL: Adversarial Prompt-based Domain Adaptation for Dialogue Summarization with Knowledge DisentanglementLulu Zhao, Fujia Zheng, Weihao Zeng, Keqing He et al.SIGIR 2022 · 6 citations
- Prompter: Zero-shot Adaptive Prefixes for Dialogue State Tracking Domain AdaptationIbrahim Taha Aksu, Min-Yen Kan, Nancy F. ChenACL 2023 · 3 citations
- Multitask Pre-training of Modular Prompt for Chinese Few-Shot LearningTianxiang Sun, Zhengfu He, Qin Zhu, Xipeng Qiu et al.ACL 2023 · 15 citations
- Less-forgetting Multi-lingual Fine-tuningYuren Mao, Yaobo Liang, Nan Duan, Haobo Wang et al.NeurIPS 2022 · 10 citations
