Sequential Reptile: Inter-Task Gradient Alignment for Multilingual Learning
Seanie Lee, Haebeom Lee, Juho Lee, Sung Ju Hwang
Abstract
Multilingual models jointly pretrained on multiple languages have achieved remarkable performance on various multilingual downstream tasks. Moreover, models finetuned on a single monolingual downstream task have shown to generalize to unseen languages. In this paper, we first show that it is crucial for those tasks to align gradients between them in order to maximize knowledge transfer while minimizing negative transfer. Despite its importance, the existing methods for gradient alignment either have a completely different purpose, ignore inter-task alignment, or aim to solve continual learning problems in rather inefficient ways. As a result of the misaligned gradients between tasks, the model suffers from severe negative transfer in the form of catastrophic forgetting of the knowledge acquired from the pretraining. To overcome the limitations, we propose a simple yet effective method that can efficiently align gradients between tasks. Specifically, we perform each inner-optimization by sequentially sampling batches from all the tasks, followed by a Reptile outer update. Thanks to the gradients aligned between tasks by our method, the model becomes less vulnerable to negative transfer and catastrophic forgetting. We extensively validate our method on various multi-task learning and zero-shot cross-lingual transfer tasks, where our method largely outperforms all the relevant baselines we consider.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f1684b1e-2708-4ac0-813e-de7909c678d2Cited by top-tier papers6
- Forget Forgetting: Continual Learning in a World of Abundant MemoryDongkyu Cho, Taesup Moon, Rumi Chunara, Kyunghyun Cho et al.ICLR 2026 · 9 citations
- Implicit meta-learning may lead language models to trust more reliable sourcesDmitrii Krasheninnikov, Egor Krasheninnikov, Bruno Kacper Mlodozeniec, Tegan Maharaj et al.ICML 2024 · 8 citations
- SeCom: On Memory Construction and Retrieval for Personalized Conversational AgentsZhuoshi Pan, Qianhui Wu, Huiqiang Jiang, Xufang Luo et al.ICLR 2025
- Analyzing and Reducing the Performance Gap in Cross-Lingual Transfer with Fine-tuning Slow and FastYiduo Guo, Yaobo Liang, Dongyan Zhao, Bing Liu et al.ACL 2023
- Balanced Direction from Multifarious Choices: Arithmetic Meta-Learning for Domain GeneralizationXiran Wang, Jian Zhang, Lei Qi, Yinghuan ShiCVPR 2025
Builds on16
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine et al.NeurIPS 2020 · 2,261 citations
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- Gradient Vaccine: Investigating and Improving Multi-task Optimization in Massively Multilingual ModelsZirui Wang, Yulia Tsvetkov, Orhan Firat, Yuan CaoICLR 2021 · 241 citations
- On the Origin of Implicit Regularization in Stochastic Gradient DescentSamuel L. Smith, Benoit Dherin, David G. T. Barrett, Soham DeICLR 2021 · 235 citations
- Multilingual Alignment of Contextual Word RepresentationsSteven Cao, Nikita Kitaev, Dan KleinICLR 2020 · 211 citations
Related papers
- Continual Learning with Global AlignmentXueying Bai, Jinghuan Shang, Yifan Sun, Niranjan BalasubramanianNeurIPS 2024 · 1 citation
- Embracing Language Inclusivity and Diversity in CLIP through Continual Language LearningBang Yang, Yong Dai, Xuxin Cheng, Yaowei Li et al.AAAI 2024 · 9 citations
- Towards Dynamic Modality Alignment in Multimodal Continual LearningJiayao Tan, Fan Lyu, Tianle Liu, Fuyuan Hu et al.CVPR 2026
- Cross-Lingual Transfer with Class-Weighted Language-Invariant RepresentationsRuicheng Xian, Heng Ji, Han ZhaoICLR 2022 · 5 citations
- Layerwise Optimization by Gradient Decomposition for Continual LearningShixiang Tang, Dapeng Chen, Jinguo Zhu, Shijie Yu et al.CVPR 2021
