UNITE: Universal kNowledge Integration from Task-specific Experts
Shuxia Lin, Qiufeng Wang 00002, Xu Yang, Xin Geng
Abstract
Large language models (LLMs) with Mixture-of-Experts (MoE) architectures achieve strong performance under sparse activation. However, their expertise is often fragmented across experts and redundant across layers. Prior studies primarily diagnosed redundancy or parameter importance, revealing overlaps but lacking mechanisms to transform them into reusable knowledge. In contrast, human learning succeeds not by memorizing isolated facts but by reusing shared strategies across domains, which motivates the question: do MoE models similarly encode universal knowledge that can be systematically extracted and reused? We propose Universal kNowledge Integration from Task-specific Experts (UNITE), a framework that consolidates experts through Fisher-weighted fusion and then applies Tucker decomposition to disentangle shared low-rank input/output subspaces as universal knowledge from layer-specific variations. This universal component provides a compact basis for reconstructing target models with flexible depth, enabling lightweight yet competitive adaptation across tasks. To assess effectiveness, we evaluate data efficiency, convergence speed, and generalization across multiple MoE-based LLMs and diverse datasets. The results show that UNITE not only extracts universal knowledge, but also flexibly enabling once-for-all extraction and flexible target model construction that generalize across domains.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- FedPAT: Federated Test-Time Adaptation via Prototype Affinity TopologyShunxin Guo, JIAQI LYU, Zhiqiang Kou, Shuxia Lin et al.ICML 2026
- DynaMem: Consistent Long Video Generation via Hierarchical Memory and Motion PriorsJingyu Lin, Xinyi Shang, Peng Sun, Cunjian Chen et al.ICML 2026
Builds on20
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- GLaM: Efficient Scaling of Language Models with Mixture-of-ExpertsNan Du, Yanping Huang, Andrew M. Dai, Simon Tong et al.ICML 2022 · 1,173 citations
- Movement Pruning: Adaptive Sparsity by Fine-TuningVictor Sanh, Thomas Wolf, Alexander M. RushNeurIPS 2020 · 656 citations
Related papers
- XPERT: Expert Knowledge Transfer for Effective Training of Language ModelsChang Liu, boyu shi, Xu Yang, Xin GengICML 2026 · 2 citations
- Each Rank Could be an Expert: Single-Ranked Mixture of Experts LoRA for Multi-task LearningZiyu Zhao, Yixiao Zhou, Xin Yu, Zhi Zhang et al.KDD 2026 · 13 citations
- MoEUT: Mixture-of-Experts Universal TransformersRóbert Csordás, Kazuki Irie, Jürgen Schmidhuber, Christopher Potts et al.NeurIPS 2024 · 65 citations
- TD-MoE: Tensor Decomposition for MoE ModelsYuebin XU, YANHONG WANG, Xuemei Peng, Hui Zang et al.ICLR 2026
- Capability Decomposition for Unified Information Extraction via Hierarchical Mixture-of-ExpertsJing Zhou, Peng Wang, Wenjun Ke, Jiajun Liu et al.ACL 2026
