Breaking the Scale Barrier: One-Shot Knowledge Transfer via Frequency Transform
Jianlu Shen, Fu Feng, Yucheng Xie, JIAQI LYU, Xin Geng
摘要
Transferring knowledge by fine-tuning large-scale pre-trained networks has become a standard paradigm for downstream tasks, yet the knowledge of a pre-trained model is tightly coupled with monolithic architecture, which restricts flexible reuse across models of varying scales. In response to this challenge, recent approaches typically resort to either parameter selection, which fails to capture the interdependent structure of this knowledge, or parameter prediction using generative models that depend on impractical access to large network collections. In this paper, we identify the low-frequency components of model weights as the concrete carrier of foundational, task-agnostic knowledge—its "learngene"—and validate this by demonstrating its efficient inheritance by downstream models and tasks. Based on this insight, we propose FRONT (FRequency dOmain kNowledge Transfer), a novel framework that uses the Discrete Cosine Transform (DCT) to isolate the low-frequency "learngene". This learngene can be seamlessly adapted to initialize models of arbitrary size via simple truncation or padding, a process that is entirely training-free. For enhanced performance, we propose an optional low-cost refinement process that introduces a spectral regularizer to further improve the learngene's transferability. Extensive experiments demonstrate that FRONT achieves the state-of-the-art performance, accelerates convergence by up to in vision tasks, and reduces training FLOPs by an average of 40.5% in language tasks. Code is available at https://github.com/LUcy0505/FRONT.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper22
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- An Empirical Study of Training Self-Supervised Vision TransformersXinlei Chen, Saining Xie, Kaiming HeICCV 2021 · 被引用 2,340 次
- Mimetic Initialization of Self-Attention LayersAsher Trockman, J. Zico KolterICML 2023 · 被引用 54 次
- Initializing Models with Larger OnesZhiqiu Xu, Yanjie Chen, Kirill Vishniakov, Yida Yin 等ICLR 2024 · 被引用 40 次
相关 Paper
- Inheriting Generalized Learngene for Efficient Knowledge Transfer across Multiple TasksYuankun Zu, Shiyu Xia, Xu Yang, Qiufeng Wang 等AAAI 2025
- Inheriting Generalizable Knowledge from LLMs to Diverse Vertical TasksChang Liu, boyu shi, Xu Yang, Qiufeng Wang 等ICLR 2026
- FINE: Factorizing Knowledge for Initialization of Variable-sized Diffusion ModelsYucheng Xie, Fu Feng, Ruixiao Shi, Jianlu Shen 等CVPR 2026
- A Unified Framework for Knowledge Transfer in Bidirectional Model ScalingJianlu Shen, Fu Feng, Jiaze Xu, Yucheng Xie 等CVPR 2026 · 被引用 1 次
- FAT: Frequency-Aware Pretraining for Enhanced Time-Series Representation LearningRui Cheng, Xiangfei Jia, Qing Li, Rong Xing 等KDD 2025 · 被引用 2 次
