Amalgamating Multi-Task Models with Heterogeneous Architectures
Jidapa Thadajarassiri, Walter Gerych, Xiangnan Kong, Elke A. Rundensteiner
摘要
Multi-task learning (MTL) is essential for real-world applications that handle multiple tasks simultaneously, such as selfdriving cars. MTL methods improve the performance of all tasks by utilizing information across tasks to learn a robust shared representation. However, acquiring sufficient labeled data tends to be extremely expensive, especially when having to support many tasks. Recently, Knowledge Amalgamation (KA) has emerged as an effective strategy for addressing the lack of labels by instead learning directly from pretrained models (teachers). KA learns one unified multi-task student that masters all tasks across all teachers. Existing KA for MTL works are limited to teachers with identical architectures, and thus propose layer-to-layer based approaches. Unfortunately, in practice, teachers may have heterogeneous architectures; their layers may not be aligned and their dimensionalities or scales may be incompatible. Amalgamating multi-task teachers with heterogeneous architectures remains an open problem. For this, we design Versatile Common Feature Consolidator (VENUS), the first solution to this problem. VENUS fuses knowledge from the shared representations of each teacher into one unified generalized representation for all tasks. Specifically, we design the Feature Consolidator network that leverages an array of teacher-specific trainable adaptors. These adaptors enable the student to learn from multiple teachers, even if they have incompatible learned representations. We demonstrate that VENUS outperforms five alternative methods on numerous benchmark datasets across a broad spectrum of experiments.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- Multi-Task Self-Training for Learning General RepresentationsGolnaz Ghiasi, Barret Zoph, Ekin D. Cubuk, Quoc V. Le 等ICCV 2021 · 被引用 119 次
- Customizing Student Networks From Heterogeneous Teachers via Adaptive Knowledge AmalgamationChengchao Shen, Mengqi Xue, Xinchao Wang, Jie Song 等ICCV 2019 · 被引用 63 次
- Semi-Supervised Knowledge Amalgamation for Sequence ClassificationJidapa Thadajarassiri, Thomas Hartvigsen, Xiangnan Kong, Elke A. RundensteinerAAAI 2021 · 被引用 14 次
- Knowledge Amalgamation for Multi-Label Classification via Label Dependency TransferJidapa Thadajarassiri, Thomas Hartvigsen, Walter Gerych, Xiangnan Kong 等AAAI 2023 · 被引用 9 次
相关 Paper
- Data-Free Knowledge Amalgamation via Group-Stack Dual-GANJingwen Ye, Yixin Ji, Xinchao Wang, Xin Gao 等CVPR 2020
- BoKA: Bayesian Optimization based Knowledge Amalgamation for Multi-unknown-domain Text ClassificationLinzhu Yu, Huan Li, Ke Chen, Lidan ShouKDD 2024 · 被引用 2 次
- Task Prototype-Based Knowledge Retrieval for Multi-Task Learning from Partially Annotated DataYoungmin Oh, Hyung-Il Kim, Jung Uk KimAAAI 2026
- Swiss Army Knife: Synergizing Biases in Knowledge from Vision Foundation Models for Multi-Task LearningYuxiang Lu, Shengcao Cao, Yu-Xiong WangICLR 2025
- MTL-KD: Multi-Task Learning Via Knowledge Distillation for Generalizable Neural Vehicle Routing SolverYuepeng Zheng, Fu Luo, Zhenkun Wang, Yaoxin Wu 等NeurIPS 2025 · 被引用 13 次
