FINE: Factorizing Knowledge for Initialization of Variable-sized Diffusion Models
Yucheng Xie, Fu Feng, Ruixiao Shi, Jianlu Shen, Jing Wang, Yong Rui, Xin Geng
Abstract
The training of diffusion models is computationally intensive, making effective pre-training essential. However, realworld deployments often demand models of variable sizes due to diverse memory and computational constraints, posing challenges when corresponding pre-trained versions are unavailable. To address this, we propose FINE, a novel pre-training method whose resulting model can flexibly factorize its knowledge into fundamental components, termed learngenes, enabling direct initialization of models of various sizes and eliminating the need for repeated pre-training. Rather than optimizing a conventional full-parameter model, FINE represents each layer's weights as the product of U ⋆ , Σ (l) ⋆ , and V ⊤ ⋆ , where U ⋆ and V ⋆ serve as size-agnostic learngenes shared across layers, while Σ (l) ⋆ remains layer-specific. By jointly training these components, FINE forms a decomposable and transferable knowledge structure that allows efficient initialization through flexible recombination of learngenes, requiring only light retraining of Σ (l)
⋆ on limited data. Extensive experiments demonstrate the efficiency of FINE, achieving state-of-the-art performance in initializing variable-sized models across diverse resource-constrained deployments. Furthermore, models initialized by FINE effectively adapt to diverse tasks, showcasing the task-agnostic versatility of learngenes.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 20f7ea9f-7961-4160-8665-37b26fd206a3Cited by top-tier papers3
- Initializing Variable-sized Vision Transformers from Learngene with Learnable TransformationShiyu Xia, Yuankun Zu, Xu Yang, Xin GengNeurIPS 2024 · 9 citations
- ECO: Evolving Core Knowledge for Efficient TransferFu Feng, Yucheng Xie, Ruixiao Shi, Jianlu Shen et al.NeurIPS 2025 · 4 citations
- A Unified Framework for Knowledge Transfer in Bidirectional Model ScalingJianlu Shen, Fu Feng, Jiaze Xu, Yucheng Xie et al.CVPR 2026 · 1 citation
Builds on37
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
Related papers
- KIND: Knowledge Integration and Diversion for Training Decomposable ModelsYucheng Xie, Fu Feng, Ruixiao Shi, Jing Wang et al.ICML 2025
- WAVE: Weight Templates for Adaptive Initialization of Variable-sized ModelsFu Feng, Yucheng Xie, Jing Wang, Xin GengCVPR 2025
- Linearly Decomposing and Recomposing Vision Transformers for Diverse-Scale ModelsShuxia Lin, Miaosen Zhang, Ruiming Chen, Xu Yang et al.NeurIPS 2024 · 7 citations
- Building Variable-Sized Models via Learngene PoolBoyu Shi, Shiyu Xia, Xu Yang, Haokun Chen et al.AAAI 2024 · 5 citations
- Breaking the Scale Barrier: One-Shot Knowledge Transfer via Frequency TransformJianlu Shen, Fu Feng, Yucheng Xie, JIAQI LYU et al.ICML 2026
