Composable Sparse Fine-Tuning for Cross-Lingual Transfer
Alan Ansell, Edoardo Maria Ponti, Anna Korhonen, Ivan Vulic
Abstract
Fine-tuning the entire set of parameters of a large pretrained model has become the mainstream approach for transfer learning. To increase its efficiency and prevent catastrophic forgetting and interference, techniques like adapters and sparse fine-tuning have been developed. Adapters are modular, as they can be combined to adapt a model towards different facets of knowledge (e.g., dedicated language and/or task adapters). Sparse finetuning is expressive, as it controls the behavior of all model components. In this work, we introduce a new fine-tuning method with both these desirable properties. In particular, we learn sparse, real-valued masks based on a simple variant of the Lottery Ticket Hypothesis. Task-specific masks are obtained from annotated data in a source language, and languagespecific masks from masked language modeling in a target language. Both these masks can then be composed with the pretrained model. Unlike adapter-based fine-tuning, this method neither increases the number of parameters at inference time nor alters the original model architecture. Most importantly, it outperforms adapters in zero-shot cross-lingual transfer by a large margin in a series of multilingual benchmarks, including Universal Dependencies, MasakhaNER, and AmericasNLI. Based on an in-depth analysis, we additionally find that sparsity is crucial to prevent both 1) interference between the fine-tunings to be composed and 2) overfitting. We release the code and models at https://github.com/ cambridgeltl/composable-sft .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 787171f5-b76a-4d43-a6b5-7e24f2acb7c8Cited by top-tier papers53
- PiSSA: Principal Singular Values and Singular Vectors Adaptation of Large Language ModelsFanxu Meng, Zhaohui Wang, Muhan ZhangNeurIPS 2024 · 374 citations
- On the Effectiveness of Parameter-Efficient Fine-TuningZihao Fu, Haoran Yang, Anthony Man-Cho So, Wai Lam et al.AAAI 2023 · 234 citations
- Parameter-Efficient Orthogonal Finetuning via Butterfly FactorizationWeiyang Liu, Zeju Qiu, Yao Feng, Yuliang Xiu et al.ICLR 2024 · 111 citations
- Model Merging by Uncertainty-Based Gradient MatchingNico Daheim, Thomas Möllenhoff, Edoardo M. Ponti, Iryna Gurevych et al.ICLR 2024 · 86 citations
- IGLUE: A Benchmark for Transfer Learning across Modalities, Tasks, and LanguagesEmanuele Bugliarello, Fangyu Liu, Jonas Pfeiffer, Siva Reddy et al.ICML 2022 · 71 citations
Builds on14
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- Comparing Rewinding and Fine-tuning in Neural Network PruningAlex Renda, Jonathan Frankle, Michael CarbinICLR 2020 · 437 citations
- Supermasks in SuperpositionMitchell Wortsman, Vivek Ramanujan, Rosanne Liu, Aniruddha Kembhavi et al.NeurIPS 2020 · 364 citations
- Proving the Lottery Ticket Hypothesis: Pruning is All You NeedEran Malach, Gilad Yehudai, Shai Shalev-Shwartz, Ohad ShamirICML 2020 · 327 citations
- Playing the lottery with rewards and multiple languages: lottery tickets in RL and NLPHaonan Yu, Sergey Edunov, Yuandong Tian, Ari S. MorcosICLR 2020 · 156 citations
Related papers
- Parameter-Efficient Fine-Tuning without Introducing New LatencyBaohao Liao, Yan Meng, Christof MonzACL 2023 · 26 citations
- Less-forgetting Multi-lingual Fine-tuningYuren Mao, Yaobo Liang, Nan Duan, Haobo Wang et al.NeurIPS 2022 · 10 citations
- Efficient Unseen Language Adaptation for Multilingual Pre-Trained Language ModelsPo-Heng Chen, Yun-Nung ChenEMNLP 2024
- Regularized Mask Tuning: Uncovering Hidden Knowledge in Pre-trained Vision-Language ModelsKecheng Zheng, Wei Wu, Ruili Feng, Kai Zhu et al.ICCV 2023 · 13 citations
- Overcoming Catastrophic Forgetting in Zero-Shot Cross-Lingual GenerationTu Vu, Aditya Barua, Brian Lester, Daniel Cer et al.EMNLP 2022 · 18 citations
