Two-Stream Transformer for Multi-Label Image Classification
Xuelin Zhu, Jiuxin Cao, Jiawei Ge, Weijia Liu, Bo Liu
摘要
Multi-label image classification is a fundamental yet challenging task in computer vision that aims to identify multiple objects from a given image. Recent studies on this task mainly focus on learning cross-modal interactions between label semantics and high-level visual representations via an attention operation. However, these one-shot attention based approaches generally perform poorly in establishing accurate and robust alignments between vision and text due to the acknowledged semantic gap. In this paper, we propose a two-stream transformer (TSFormer) learning framework, in which the spatial stream focuses on extracting patch features with a global perception, while the semantic stream aims to learn vision-aware label semantics as well as their correlations via a multi-shot attention mechanism. Specifically, in each layer of TSFormer, a cross-modal attention module is developed to aggregate visual features from spatial stream into semantic stream and update label semantics via a residual connection. In this way, the semantic gap between two streams gradually narrows as the procedure progresses layer by layer, allowing the semantic stream to produce sophisticated visual representations for each label towards accurate label recognition. Extensive experiments on three visual benchmarks, including Pascal VOC 2007, Microsoft COCO and NUS-WIDE, consistently demonstrate that our proposed TSFormer achieves state-of-the-art performance on the multi-label image classification task.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper5
- Scene-Aware Label Graph Learning for Multi-Label Image ClassificationXuelin Zhu, Jian Liu, Weijia Liu, Jiawei Ge 等ICCV 2023 · 被引用 39 次
- PatchCT: Aligning Patch Set and Label Set with Conditional Transport for Multi-Label Image ClassificationMiaoge Li, Dongsheng Wang, Xinyang Liu, Zequn Zeng 等ICCV 2023 · 被引用 30 次
- Specifying What You Know or Not for Multi-Label Class-Incremental LearningAoting Zhang, Dongbao Yang, Chang Liu, Xiaopeng Hong 等AAAI 2025 · 被引用 6 次
- HSVLT: Hierarchical Scale-Aware Vision-Language Transformer for Multi-Label Image ClassificationShuyi Ouyang, Hongyi Wang, Ziwei Niu, Zhenjia Bai 等ACM MM 2023 · 被引用 5 次
- MambaMl: Exploring State Space Models for Multi-Label Image ClassificationXuelin Zhu, Jian Liu, Jiuxin Cao, Bing WangICCV 2025 · 被引用 2 次
相关 Paper
- M3TR: Multi-modal Multi-label Recognition with TransformerJiawei Zhao, Yifan Zhao, Jia LiACM MM 2021 · 被引用 45 次
- Unifying Two-Stream Encoders with Transformers for Cross-Modal RetrievalYi Bin, Haoxuan Li, Yahui Xu, Xing Xu 等ACM MM 2023 · 被引用 33 次
- Cross-Modality Attention with Semantic Graph Embedding for Multi-Label ClassificationRenchun You, Zhiyao Guo, Lei Cui, Xiang Long 等AAAI 2020 · 被引用 221 次
- Improving Image Captioning through Visual and Semantic Mutual PromotionJing Zhang, Yingshuai Xie, Xiaoqiang LiuACM MM 2023 · 被引用 4 次
- Transformer-based Dual Relation Graph for Multi-label Image RecognitionJiawei Zhao, Ke Yan, Yifan Zhao, Xiaowei Guo 等ICCV 2021 · 被引用 109 次
