MetaPerturb: Transferable Regularizer for Heterogeneous Tasks and Architectures
Jeongun Ryu, Jaewoong Shin, Haebeom Lee, Sung Ju Hwang
摘要
Regularization and transfer learning are two popular techniques to enhance model generalization on unseen data, which is a fundamental problem of machine learning. Regularization techniques are versatile, as they are task-and architecture-agnostic, but they do not exploit a large amount of data available. Transfer learning methods learn to transfer knowledge from one domain to another, but may not generalize across tasks and architectures, and may introduce new training cost for adapting to the target task. To bridge the gap between the two, we propose a transferable perturbation, MetaPerturb, which is meta-learned to improve generalization performance on unseen data. MetaPerturb is implemented as a set-based lightweight network that is agnostic to the size and the order of the input, which is shared across the layers. Then, we propose a meta-learning framework, to jointly train the perturbation function over heterogeneous tasks in parallel. As MetaPerturb is a set-function trained over diverse distributions across layers and tasks, it can generalize to heterogeneous tasks and architectures. We validate the efficacy and generality of MetaPerturb trained on a specific source domain and architecture, by applying it to the training of diverse neural architectures on heterogeneous target datasets against various regularizers and fine-tuning. The results show that the networks trained with MetaPerturb significantly outperform the baselines on most of the tasks and architectures, with a negligible increase in the parameter size and no hyperparameters to tune.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Learning Cross-Image Object Semantic Relation in Transformer for Few-Shot Fine-Grained Image ClassificationBo Zhang, Jiakang Yuan, Baopu Li, Tao Chen 等ACM MM 2022 · 被引用 42 次
- On the Importance of Distractors for Few-Shot ClassificationRajshekhar Das, Yu-Xiong Wang, José M. F. MouraICCV 2021 · 被引用 36 次
- Auto-Transfer: Learning to Route Transferable RepresentationsKeerthiram Murugesan, Vijay Sadashivaiah, Ronny Luss, Karthikeyan Shanmugam 等ICLR 2022 · 被引用 6 次
- Online Hyperparameter Meta-Learning with Hypergradient DistillationHaebeom Lee, Hayeon Lee, Jaewoong Shin, Eunho Yang 等ICLR 2022 · 被引用 6 次
- Latent Paraphrasing: Perturbation on Layers Improves Knowledge Injection in Language ModelsMinki Kang, Sung Ju Hwang, Gibbeum Lee, Jaewoong ChoNeurIPS 2024 · 被引用 3 次
它引用的顶会 Paper1
相关 Paper
- Meta Distant Transfer Learning for Pre-trained Language ModelsChengyu Wang, Haojie Pan, Minghui Qiu, Jun Huang 等EMNLP 2021 · 被引用 3 次
- Meta-Learning with Fewer Tasks through Task InterpolationHuaxiu Yao, Linjun Zhang, Chelsea FinnICLR 2022 · 被引用 66 次
- Open Domain Generalization with Domain-Augmented Meta-LearningYang Shu, Zhangjie Cao, Chenyu Wang, Jianmin Wang 等CVPR 2021
- Improving Generalization of Meta-Learning with Inverted Regularization at Inner-LevelLianzhe Wang, Shiji Zhou, Shanghang Zhang, Xu Chu 等CVPR 2023
- Set-based Meta-Interpolation for Few-Task Meta-LearningSeanie Lee, Bruno Andreis, Kenji Kawaguchi, Juho Lee 等NeurIPS 2022 · 被引用 13 次
