Hub-Pathway: Transfer Learning from A Hub of Pre-trained Models
Yang Shu, Zhangjie Cao, Ziyang Zhang, Jianmin Wang, Mingsheng Long
Abstract
Transfer learning aims to leverage knowledge from pre-trained models to benefit the target task. Prior transfer learning work mainly transfers from a single model. However, with the emergence of deep models pre-trained from different resources, model hubs consisting of diverse models with various architectures, pre-trained datasets and learning paradigms are available. Directly applying single-model transfer learning methods to each model wastes the abundant knowledge of the model hub and suffers from high computational cost. In this paper, we propose a Hub-Pathway framework to enable knowledge transfer from a model hub. The framework generates data-dependent pathway weights, based on which we assign the pathway routes at the input level to decide which pre-trained models are activated and passed through, and then set the pathway aggregation at the output level to aggregate the knowledge from different models to make predictions. The proposed framework can be trained end-to-end with the target task-specific loss, where it learns to explore better pathway configurations and exploit the knowledge in pre-trained models for each target datum. We utilize a noisy pathway generator and design an exploration loss to further explore different pathways throughout the model hub. To fully exploit the knowledge in pre-trained models, each model is further trained by specific data that activate it, which ensures its performance and enhances knowledge transfer. Experiment results on computer vision and reinforcement learning tasks demonstrate that the proposed Hub-Pathway framework achieves the state-of-the-art performance for model hub transfer learning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f4241626-48fe-4dfa-9f49-cf9bfe10ac30Cited by top-tier papers4
- Quick-Tune: Quickly Learning Which Pretrained Model to Finetune and HowSebastian Pineda-Arango, Fabio Ferreira, Arlind Kadra, Frank Hutter et al.ICLR 2024 · 27 citations
- Agile Multi-Source-Free Domain AdaptationXinyao Li, Jingjing Li, Fengling Li, Lei Zhu et al.AAAI 2024 · 23 citations
- ToVE: Efficient Vision-Language Learning via Knowledge Transfer from Vision ExpertsYuanchen Wu, Junlong Du, Ke Yan, Shouhong Ding et al.ICLR 2025
- Efficient Stitchable Task AdaptationHaoyu He, Zizheng Pan, Jing Liu, Jianfei Cai et al.CVPR 2024
Builds on20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le et al.ICCV 2019 · 9,163 citations
Related papers
- Zoo-Tuning: Adaptive Transfer from A Zoo of ModelsYang Shu, Zhi Kou, Zhangjie Cao, Jianmin Wang et al.ICML 2021 · 46 citations
- Which Model to Transfer? Finding the Needle in the Growing HaystackCédric Renggli, André Susano Pinto, Luka Rimanic, Joan Puigcerver et al.CVPR 2022 · 13 citations
- Harnessing Language Model for Cross-Heterogeneity Graph Knowledge TransferJinyu Yang, Ruijia Wang, Cheng Yang, Bo Yan et al.AAAI 2025 · 4 citations
- UniDense: Unleashing Diffusion Models with Meta-Routers for Universal Few-Shot Dense PredictionLintao Dong, Wei Zhai, Zheng-Jun ZhaACM MM 2024 · 1 citation
- DEPARA: Deep Attribution Graph for Deep Knowledge TransferabilityJie Song, Yixin Chen, Jingwen Ye, Xinchao Wang et al.CVPR 2020
