Hybrid Sharing for Multi-Label Image Classification
Zihao Yin, Chen Gan, Kelei He, Yang Gao, Junfeng Zhang
Abstract
Existing multi-label classification methods have long suffered from label heterogeneity, where learning a label obscures another. By modeling multi-label classification as a multi-task problem, this issue can be regarded as a negative transfer, which indicates challenges to achieve simultaneously satisfied performance across multiple tasks. In this work, we propose the Hybrid Sharing Query (HSQ), a transformer-based model that introduces the mixture-of-experts architecture to image multi-label classification. HSQ is designed to leverage label correlations while mitigating heterogeneity effectively. To this end, HSQ is incorporated with a fusion expert framework that enables it to optimally combine the strengths of task-specialized experts with shared experts, ultimately enhancing multi-label classification performance across most labels. Extensive experiments are conducted on two benchmark datasets, with the results demonstrating that the proposed method achieves state-of-the-art performance and yields simultaneous improvements across most labels. The code is available at this URL.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- MAT-Agent: Adaptive Multi-Agent Training OptimizationJusheng Zhang, Kaitong Cai, Yijia Fan, Ningyuan Liu et al.NeurIPS 2025 · 46 citations
- Correlative and Discriminative Label Grouping for Multi-Label Visual Prompt TuningLei-Lei Ma, Shuo Xu, Ming-Kun Xie, Lei Wang et al.CVPR 2025
Builds on12
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- CvT: Introducing Convolutions to Vision TransformersHaiping Wu, Bin Xiao, Noel Codella, Mengchen Liu et al.ICCV 2021 · 2,397 citations
- Mixture-of-Experts with Expert Choice RoutingYanqi Zhou, Tao Lei, Hanxiao Liu, Nan Du et al.NeurIPS 2022 · 933 citations
- Learning Semantic-Specific Graph Representation for Multi-Label Image RecognitionTianshui Chen, Muxin Xu, Xiaolu Hui, Hefeng Wu et al.ICCV 2019 · 347 citations
Related papers
- Language-Guided Transformer for Federated Multi-Label ClassificationI-Jieh Liu, Ci-Siang Lin, Fu-En Yang, Yu-Chiang Frank WangAAAI 2024 · 16 citations
- Mod-Squad: Designing Mixtures of Experts As Modular Multi-Task LearnersZitian Chen, Yikang Shen, Mingyu Ding, Zhenfang Chen et al.CVPR 2023
- View-Category Interactive Sharing Transformer for Incomplete Multi-View Multi-Label LearningShilong Ou, Zhe Xue, Yawen Li, Meiyu Liang et al.CVPR 2024 · 11 citations
- HSVLT: Hierarchical Scale-Aware Vision-Language Transformer for Multi-Label Image ClassificationShuyi Ouyang, Hongyi Wang, Ziwei Niu, Zhenjia Bai et al.ACM MM 2023 · 5 citations
- General Multi-Label Image Classification With TransformersJack Lanchantin, Tianlu Wang, Vicente Ordonez, Yanjun QiCVPR 2021
