KIND: Knowledge Integration and Diversion for Training Decomposable Models
Yucheng Xie, Fu Feng, Ruixiao Shi, Jing Wang, Yong Rui, Xin Geng
Abstract
Pre-trained models have become the preferred backbone due to the increasing complexity of model parameters. However, traditional pretrained models often face deployment challenges due to their fixed sizes, and are prone to negative transfer when discrepancies arise between training tasks and target tasks. To address this, we propose KIND, a novel pre-training method designed to construct decomposable models. KIND integrates knowledge by incorporating Singular Value Decomposition (SVD) as a structural constraint, with each basic component represented as a combination of a column vector, singular value, and row vector from U , Σ, and V ⊤ matrices. These components are categorized into learngenes for encapsulating class-agnostic knowledge and tailors for capturing class-specific knowledge, with knowledge diversion facilitated by a class gate mechanism during training. Extensive experiments demonstrate that models pretrained with KIND can be decomposed into learngenes and tailors, which can be adaptively recombined for diverse resource-constrained deployments. Moreover, for tasks with large domain shifts, transferring only learngenes with task-agnostic knowledge, when combined with randomly initialized tailors, effectively mitigates domain shifts. Code will be made available at https://github.com/Te4P0t/KIND .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- DivControl: Knowledge Diversion for Controllable Image GenerationYucheng Xie, Fu Feng, Ruixiao Shi, Jing Wang et al.AAAI 2026 · 4 citations
- A Unified Framework for Knowledge Transfer in Bidirectional Model ScalingJianlu Shen, Fu Feng, Jiaze Xu, Yucheng Xie et al.CVPR 2026 · 1 citation
- FINE: Factorizing Knowledge for Initialization of Variable-sized Diffusion ModelsYucheng Xie, Fu Feng, Ruixiao Shi, Jianlu Shen et al.CVPR 2026
- Breaking Semantic Boundaries: Distribution-Guided Semantic Exploration for Creative GenerationFu Feng, Yucheng Xie, Ruixiao Shi, Xu Yang et al.CVPR 2026
- Breaking the Scale Barrier: One-Shot Knowledge Transfer via Frequency TransformJianlu Shen, Fu Feng, Yucheng Xie, JIAQI LYU et al.ICML 2026
Builds on13
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
- AdaptFormer: Adapting Vision Transformers for Scalable Visual RecognitionShoufa Chen, Chongjian Ge, Zhan Tong, Jiangliu Wang et al.NeurIPS 2022 · 1,291 citations
Related papers
- Linearly Decomposing and Recomposing Vision Transformers for Diverse-Scale ModelsShuxia Lin, Miaosen Zhang, Ruiming Chen, Xu Yang et al.NeurIPS 2024 · 7 citations
- Inheriting Generalized Learngene for Efficient Knowledge Transfer across Multiple TasksYuankun Zu, Shiyu Xia, Xu Yang, Qiufeng Wang et al.AAAI 2025
- Low-Rank Knowledge Decomposition for Medical Foundation ModelsYuhang Zhou, Haolin Li, Siyuan Du, Jiangchao Yao et al.CVPR 2024 · 4 citations
- Knowledge Diversion for Efficient Morphology Control and Policy TransferFu Feng, Ruixiao Shi, Yucheng Xie, Jianlu Shen et al.ICML 2026 · 1 citation
- Building Variable-Sized Models via Learngene PoolBoyu Shi, Shiyu Xia, Xu Yang, Haokun Chen et al.AAAI 2024 · 5 citations
