Two-Stream Transformer for Multi-Label Image Classification
Xuelin Zhu, Jiuxin Cao, Jiawei Ge, Weijia Liu, Bo Liu
Abstract
Multi-label image classification is a fundamental yet challenging task in computer vision that aims to identify multiple objects from a given image. Recent studies on this task mainly focus on learning cross-modal interactions between label semantics and high-level visual representations via an attention operation. However, these one-shot attention based approaches generally perform poorly in establishing accurate and robust alignments between vision and text due to the acknowledged semantic gap. In this paper, we propose a two-stream transformer (TSFormer) learning framework, in which the spatial stream focuses on extracting patch features with a global perception, while the semantic stream aims to learn vision-aware label semantics as well as their correlations via a multi-shot attention mechanism. Specifically, in each layer of TSFormer, a cross-modal attention module is developed to aggregate visual features from spatial stream into semantic stream and update label semantics via a residual connection. In this way, the semantic gap between two streams gradually narrows as the procedure progresses layer by layer, allowing the semantic stream to produce sophisticated visual representations for each label towards accurate label recognition. Extensive experiments on three visual benchmarks, including Pascal VOC 2007, Microsoft COCO and NUS-WIDE, consistently demonstrate that our proposed TSFormer achieves state-of-the-art performance on the multi-label image classification task.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 94430b82-a690-4cc4-a826-8fa8c1c89827Cited by top-tier papers5
- Scene-Aware Label Graph Learning for Multi-Label Image ClassificationXuelin Zhu, Jian Liu, Weijia Liu, Jiawei Ge et al.ICCV 2023 · 39 citations
- PatchCT: Aligning Patch Set and Label Set with Conditional Transport for Multi-Label Image ClassificationMiaoge Li, Dongsheng Wang, Xinyang Liu, Zequn Zeng et al.ICCV 2023 · 30 citations
- Specifying What You Know or Not for Multi-Label Class-Incremental LearningAoting Zhang, Dongbao Yang, Chang Liu, Xiaopeng Hong et al.AAAI 2025 · 6 citations
- HSVLT: Hierarchical Scale-Aware Vision-Language Transformer for Multi-Label Image ClassificationShuyi Ouyang, Hongyi Wang, Ziwei Niu, Zhenjia Bai et al.ACM MM 2023 · 5 citations
- MambaMl: Exploring State Space Models for Multi-Label Image ClassificationXuelin Zhu, Jian Liu, Jiuxin Cao, Bing WangICCV 2025 · 2 citations
Related papers
- M3TR: Multi-modal Multi-label Recognition with TransformerJiawei Zhao, Yifan Zhao, Jia LiACM MM 2021 · 45 citations
- Unifying Two-Stream Encoders with Transformers for Cross-Modal RetrievalYi Bin, Haoxuan Li, Yahui Xu, Xing Xu et al.ACM MM 2023 · 33 citations
- Cross-Modality Attention with Semantic Graph Embedding for Multi-Label ClassificationRenchun You, Zhiyao Guo, Lei Cui, Xiang Long et al.AAAI 2020 · 221 citations
- Improving Image Captioning through Visual and Semantic Mutual PromotionJing Zhang, Yingshuai Xie, Xiaoqiang LiuACM MM 2023 · 4 citations
- Transformer-based Dual Relation Graph for Multi-label Image RecognitionJiawei Zhao, Ke Yan, Yifan Zhao, Xiaowei Guo et al.ICCV 2021 · 109 citations
