MetaPerturb: Transferable Regularizer for Heterogeneous Tasks and Architectures
Jeongun Ryu, Jaewoong Shin, Haebeom Lee, Sung Ju Hwang
Abstract
Regularization and transfer learning are two popular techniques to enhance model generalization on unseen data, which is a fundamental problem of machine learning. Regularization techniques are versatile, as they are task-and architecture-agnostic, but they do not exploit a large amount of data available. Transfer learning methods learn to transfer knowledge from one domain to another, but may not generalize across tasks and architectures, and may introduce new training cost for adapting to the target task. To bridge the gap between the two, we propose a transferable perturbation, MetaPerturb, which is meta-learned to improve generalization performance on unseen data. MetaPerturb is implemented as a set-based lightweight network that is agnostic to the size and the order of the input, which is shared across the layers. Then, we propose a meta-learning framework, to jointly train the perturbation function over heterogeneous tasks in parallel. As MetaPerturb is a set-function trained over diverse distributions across layers and tasks, it can generalize to heterogeneous tasks and architectures. We validate the efficacy and generality of MetaPerturb trained on a specific source domain and architecture, by applying it to the training of diverse neural architectures on heterogeneous target datasets against various regularizers and fine-tuning. The results show that the networks trained with MetaPerturb significantly outperform the baselines on most of the tasks and architectures, with a negligible increase in the parameter size and no hyperparameters to tune.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2200e2e6-1ef5-419b-9864-2c578fe14a51Cited by top-tier papers6
- Learning Cross-Image Object Semantic Relation in Transformer for Few-Shot Fine-Grained Image ClassificationBo Zhang, Jiakang Yuan, Baopu Li, Tao Chen et al.ACM MM 2022 · 42 citations
- On the Importance of Distractors for Few-Shot ClassificationRajshekhar Das, Yu-Xiong Wang, José M. F. MouraICCV 2021 · 36 citations
- Auto-Transfer: Learning to Route Transferable RepresentationsKeerthiram Murugesan, Vijay Sadashivaiah, Ronny Luss, Karthikeyan Shanmugam et al.ICLR 2022 · 6 citations
- Online Hyperparameter Meta-Learning with Hypergradient DistillationHaebeom Lee, Hayeon Lee, Jaewoong Shin, Eunho Yang et al.ICLR 2022 · 6 citations
- Latent Paraphrasing: Perturbation on Layers Improves Knowledge Injection in Language ModelsMinki Kang, Sung Ju Hwang, Gibbeum Lee, Jaewoong ChoNeurIPS 2024 · 3 citations
Builds on1
Related papers
- Meta Distant Transfer Learning for Pre-trained Language ModelsChengyu Wang, Haojie Pan, Minghui Qiu, Jun Huang et al.EMNLP 2021 · 3 citations
- Meta-Learning with Fewer Tasks through Task InterpolationHuaxiu Yao, Linjun Zhang, Chelsea FinnICLR 2022 · 66 citations
- Open Domain Generalization with Domain-Augmented Meta-LearningYang Shu, Zhangjie Cao, Chenyu Wang, Jianmin Wang et al.CVPR 2021
- Improving Generalization of Meta-Learning with Inverted Regularization at Inner-LevelLianzhe Wang, Shiji Zhou, Shanghang Zhang, Xu Chu et al.CVPR 2023
- Set-based Meta-Interpolation for Few-Task Meta-LearningSeanie Lee, Bruno Andreis, Kenji Kawaguchi, Juho Lee et al.NeurIPS 2022 · 13 citations
