Inheriting Generalizable Knowledge from LLMs to Diverse Vertical Tasks
Chang Liu, boyu shi, Xu Yang, Qiufeng Wang, Xin Geng
Abstract
Large language models (LLMs) have demonstrated remarkable generalization across diverse tasks, suggesting the existence of task-agnostic, generalizable knowledge encoded within them. However, how to systematically extract and evaluate this knowledge remains unexplored. In this work, we innovatively propose MASA (Matrix-level Alignment and Scalable Adaptation), a unified framework for extracting and transferring generalizable knowledge from LLMs. MASA first introduces a lightweight set of gene matrices trained with a dual alignment strategy, combining output alignment and spectral alignment, to capture the generalizable knowledge encoded in the feed-forward networks (FFNs) of LLM. It then employs scalable adaptation to flexibly reshape these gene matrices to match the parameter dimensions of lightweight dense models of various sizes, enabling direct initialization of their FFN layers. To evaluate the inherited knowledge, we measure the downstream performance of lightweight models initialized with MASA across language understanding and dialogue generation tasks spanning diverse vertical domains. Experiments on both dense and Mixture-of-Experts (MoE) source LLMs show that MASA consistently outperforms baselines such as random initialization, pruning, and distillation, yielding lightweight models that achieve stronger performance, require less pre-training data, and converge faster. These results establish MASA as an effective and general framework for extracting and leveraging the generalizable knowledge within LLMs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Energy-Structured Low-Rank Adaptation for Continual LearningLonghua Li, Lei Qi, Qi Tian, Xin GengICML 2026 · 1 citation
- FedPAT: Federated Test-Time Adaptation via Prototype Affinity TopologyShunxin Guo, JIAQI LYU, Zhiqiang Kou, Shuxia Lin et al.ICML 2026
- Breaking the Scale Barrier: One-Shot Knowledge Transfer via Frequency TransformJianlu Shen, Fu Feng, Yucheng Xie, JIAQI LYU et al.ICML 2026
Builds on25
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 3,415 citations
Related papers
- XPERT: Expert Knowledge Transfer for Effective Training of Language ModelsChang Liu, boyu shi, Xu Yang, Xin GengICML 2026 · 2 citations
- UNITE: Universal kNowledge Integration from Task-specific ExpertsShuxia Lin, Qiufeng Wang 00002, Xu Yang, Xin GengICLR 2026
- LLM DNA: Tracing Model Evolution via Functional RepresentationsZhaomin Wu, Haodong Zhao, Ziyang Wang, Jizhou Guo et al.ICLR 2026 · 21 citations
- Towards Efficient Dialogue Pre-training with Transferable and Interpretable Latent StructureXueliang Zhao, Lemao Liu, Tingchen Fu, Shuming Shi et al.EMNLP 2022 · 3 citations
- Split-Merge: Scalable and Memory-Efficient Merging of Expert LLMsSruthi Gorantla, Aditya Rawal, Devamanyu Hazarika, Kaixiang Lin et al.EMNLP 2025
