Semi-Supervised Knowledge Amalgamation for Sequence Classification
Jidapa Thadajarassiri, Thomas Hartvigsen, Xiangnan Kong, Elke A. Rundensteiner
Abstract
Sequence classification is essential for domains from medical diagnosis to online advertising. In these settings, data are typically proprietary, and annotations are expensive to acquire. Often times, so few annotations are available that training a robust model from scratch is impractical. Recently, knowledge amalgamation (KA) has emerged as a promising strategy for training models without this hard-to-come-by labeled training dataset. To achieve this, KA methods combine the knowledge of multiple pre-trained teacher models (trained on different classification tasks and proprietary datasets) into one student model that becomes an expert on the union of all teachers’ classes. However, we demonstrate that the state-of-the-art solutions fail in the presence of overconfident teachers, which make confident but incorrect predictions for instances from classes upon which they were not trained. Additionally, to-date no work has explored KA for sequence models. Therefore, we propose and then solve the open problem of semi-supervised KA for sequence classification (SKA). Our SKA approach first learns to estimate how trustworthy each teacher is for a given instance, then rescales the predicted probabilities from all teachers to supervise a student model. Our solution overcomes overconfident teachers through careful use of a very small amount of labeled instances. We demonstrate that this approach beats eight state-of-the-art alternatives on four real-world datasets by on average 15% in accuracy with as little as 2% of training data being annotated.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 29a72e2a-195c-49d7-a0d3-9b99336b5627Cited by top-tier papers3
- Training-Free Pretrained Model MergingZhengqi Xu, Ke Yuan, Huiqiong Wang, Yong Wang et al.CVPR 2024 · 6 citations
- Amalgamating Multi-Task Models with Heterogeneous ArchitecturesJidapa Thadajarassiri, Walter Gerych, Xiangnan Kong, Elke A. RundensteinerAAAI 2024 · 1 citation
- Swiss Army Knife: Synergizing Biases in Knowledge from Vision Foundation Models for Multi-Task LearningYuxiang Lu, Shengcao Cao, Yu-Xiong WangICLR 2025
Builds on3
- Customizing Student Networks From Heterogeneous Teachers via Adaptive Knowledge AmalgamationChengchao Shen, Mengqi Xue, Xinchao Wang, Jie Song et al.ICCV 2019 · 63 citations
- Recurrent Halting Chain for Early Multi-label ClassificationThomas Hartvigsen, Cansu Sen, Xiangnan Kong, Elke A. RundensteinerKDD 2020 · 18 citations
- Instance-Wise Dynamic Sensor Selection for Human Activity RecognitionXiaodong Yang, Yiqiang Chen, Hanchao Yu, Yingwei Zhang et al.AAAI 2020 · 10 citations
Related papers
- BoKA: Bayesian Optimization based Knowledge Amalgamation for Multi-unknown-domain Text ClassificationLinzhu Yu, Huan Li, Ke Chen, Lidan ShouKDD 2024 · 2 citations
- Knowledge Amalgamation for Multi-Label Classification via Label Dependency TransferJidapa Thadajarassiri, Thomas Hartvigsen, Walter Gerych, Xiangnan Kong et al.AAAI 2023 · 9 citations
- Low Resource Sequence Tagging with Weak LabelsEdwin Simpson, Jonas Pfeiffer, Iryna GurevychAAAI 2020 · 14 citations
- Data-Free Knowledge Amalgamation via Group-Stack Dual-GANJingwen Ye, Yixin Ji, Xinchao Wang, Xin Gao et al.CVPR 2020
- Amalgamating Knowledge From Heterogeneous Graph Neural NetworksYongcheng Jing, Yiding Yang, Xinchao Wang, Mingli Song et al.CVPR 2021
