Amalgamating Multi-Task Models with Heterogeneous Architectures
Jidapa Thadajarassiri, Walter Gerych, Xiangnan Kong, Elke A. Rundensteiner
Abstract
Multi-task learning (MTL) is essential for real-world applications that handle multiple tasks simultaneously, such as selfdriving cars. MTL methods improve the performance of all tasks by utilizing information across tasks to learn a robust shared representation. However, acquiring sufficient labeled data tends to be extremely expensive, especially when having to support many tasks. Recently, Knowledge Amalgamation (KA) has emerged as an effective strategy for addressing the lack of labels by instead learning directly from pretrained models (teachers). KA learns one unified multi-task student that masters all tasks across all teachers. Existing KA for MTL works are limited to teachers with identical architectures, and thus propose layer-to-layer based approaches. Unfortunately, in practice, teachers may have heterogeneous architectures; their layers may not be aligned and their dimensionalities or scales may be incompatible. Amalgamating multi-task teachers with heterogeneous architectures remains an open problem. For this, we design Versatile Common Feature Consolidator (VENUS), the first solution to this problem. VENUS fuses knowledge from the shared representations of each teacher into one unified generalized representation for all tasks. Specifically, we design the Feature Consolidator network that leverages an array of teacher-specific trainable adaptors. These adaptors enable the student to learn from multiple teachers, even if they have incompatible learned representations. We demonstrate that VENUS outperforms five alternative methods on numerous benchmark datasets across a broad spectrum of experiments.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 344f1122-4e03-4d1d-9d23-666ca2dae3ccBuilds on4
- Multi-Task Self-Training for Learning General RepresentationsGolnaz Ghiasi, Barret Zoph, Ekin D. Cubuk, Quoc V. Le et al.ICCV 2021 · 119 citations
- Customizing Student Networks From Heterogeneous Teachers via Adaptive Knowledge AmalgamationChengchao Shen, Mengqi Xue, Xinchao Wang, Jie Song et al.ICCV 2019 · 63 citations
- Semi-Supervised Knowledge Amalgamation for Sequence ClassificationJidapa Thadajarassiri, Thomas Hartvigsen, Xiangnan Kong, Elke A. RundensteinerAAAI 2021 · 14 citations
- Knowledge Amalgamation for Multi-Label Classification via Label Dependency TransferJidapa Thadajarassiri, Thomas Hartvigsen, Walter Gerych, Xiangnan Kong et al.AAAI 2023 · 9 citations
Related papers
- Data-Free Knowledge Amalgamation via Group-Stack Dual-GANJingwen Ye, Yixin Ji, Xinchao Wang, Xin Gao et al.CVPR 2020
- BoKA: Bayesian Optimization based Knowledge Amalgamation for Multi-unknown-domain Text ClassificationLinzhu Yu, Huan Li, Ke Chen, Lidan ShouKDD 2024 · 2 citations
- Task Prototype-Based Knowledge Retrieval for Multi-Task Learning from Partially Annotated DataYoungmin Oh, Hyung-Il Kim, Jung Uk KimAAAI 2026
- Swiss Army Knife: Synergizing Biases in Knowledge from Vision Foundation Models for Multi-Task LearningYuxiang Lu, Shengcao Cao, Yu-Xiong WangICLR 2025
- MTL-KD: Multi-Task Learning Via Knowledge Distillation for Generalizable Neural Vehicle Routing SolverYuepeng Zheng, Fu Luo, Zhenkun Wang, Yaoxin Wu et al.NeurIPS 2025 · 13 citations
