Hub-Pathway: Transfer Learning from A Hub of Pre-trained Models
Yang Shu, Zhangjie Cao, Ziyang Zhang, Jianmin Wang, Mingsheng Long
摘要
Transfer learning aims to leverage knowledge from pre-trained models to benefit the target task. Prior transfer learning work mainly transfers from a single model. However, with the emergence of deep models pre-trained from different resources, model hubs consisting of diverse models with various architectures, pre-trained datasets and learning paradigms are available. Directly applying single-model transfer learning methods to each model wastes the abundant knowledge of the model hub and suffers from high computational cost. In this paper, we propose a Hub-Pathway framework to enable knowledge transfer from a model hub. The framework generates data-dependent pathway weights, based on which we assign the pathway routes at the input level to decide which pre-trained models are activated and passed through, and then set the pathway aggregation at the output level to aggregate the knowledge from different models to make predictions. The proposed framework can be trained end-to-end with the target task-specific loss, where it learns to explore better pathway configurations and exploit the knowledge in pre-trained models for each target datum. We utilize a noisy pathway generator and design an exploration loss to further explore different pathways throughout the model hub. To fully exploit the knowledge in pre-trained models, each model is further trained by specific data that activate it, which ensures its performance and enhances knowledge transfer. Experiment results on computer vision and reinforcement learning tasks demonstrate that the proposed Hub-Pathway framework achieves the state-of-the-art performance for model hub transfer learning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Quick-Tune: Quickly Learning Which Pretrained Model to Finetune and HowSebastian Pineda-Arango, Fabio Ferreira, Arlind Kadra, Frank Hutter 等ICLR 2024 · 被引用 27 次
- Agile Multi-Source-Free Domain AdaptationXinyao Li, Jingjing Li, Fengling Li, Lei Zhu 等AAAI 2024 · 被引用 23 次
- ToVE: Efficient Vision-Language Learning via Knowledge Transfer from Vision ExpertsYuanchen Wu, Junlong Du, Ke Yan, Shouhong Ding 等ICLR 2025
- Efficient Stitchable Task AdaptationHaoyu He, Zizheng Pan, Jing Liu, Jianfei Cai 等CVPR 2024
它引用的顶会 Paper20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le 等ICCV 2019 · 被引用 9,163 次
相关 Paper
- Zoo-Tuning: Adaptive Transfer from A Zoo of ModelsYang Shu, Zhi Kou, Zhangjie Cao, Jianmin Wang 等ICML 2021 · 被引用 46 次
- Which Model to Transfer? Finding the Needle in the Growing HaystackCédric Renggli, André Susano Pinto, Luka Rimanic, Joan Puigcerver 等CVPR 2022 · 被引用 13 次
- Harnessing Language Model for Cross-Heterogeneity Graph Knowledge TransferJinyu Yang, Ruijia Wang, Cheng Yang, Bo Yan 等AAAI 2025 · 被引用 4 次
- UniDense: Unleashing Diffusion Models with Meta-Routers for Universal Few-Shot Dense PredictionLintao Dong, Wei Zhai, Zheng-Jun ZhaACM MM 2024 · 被引用 1 次
- DEPARA: Deep Attribution Graph for Deep Knowledge TransferabilityJie Song, Yixin Chen, Jingwen Ye, Xinchao Wang 等CVPR 2020
