M3TR: Multi-modal Multi-label Recognition with Transformer
Jiawei Zhao, Yifan Zhao, Jia Li
Abstract
Multi-label image recognition aims to recognize multiple objects simultaneously in one image. Recent ideas to solve this problem have focused on learning dependencies of label co-occurrences to enhance the high-level semantic representations. However, these methods usually neglect the important relations of intrinsic visual structures and face difficulties in understanding contextual relationships. To build the global scope of visual context as well as interactions between visual modality and linguistic modality, we propose the Multi-Modal Multi-label recognition TRansformers (M3TR) with the ternary relationship learning for inter-and intra-modalities. For the intra-modal relationship, we make insightful conjunction of CNNs and Transformers, which embeds visual structures into the high-level features by learning the semantic cross-attention. For constructing the interactions between the visual and linguistic modalities, we propose a linguistic cross-attention to embed the class-wise linguistic information into the visual structure learning, and finally present a linguistic guided enhancement module to enhance the representation of high-level semantics. Experimental evidence reveals that with the collaborative learning of ternary relationship, our proposed M3TR achieves new state-of-the-art on two public multi-label recognition benchmarks.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 8a026774-c67b-4618-b8f1-9371f196b216Cited by top-tier papers5
- PatchCT: Aligning Patch Set and Label Set with Conditional Transport for Multi-Label Image ClassificationMiaoge Li, Dongsheng Wang, Xinyang Liu, Zequn Zeng et al.ICCV 2023 · 30 citations
- Multi-modal Extreme ClassificationAnshul Mittal, Kunal Dahiya, Shreya Malani, Janani Ramaswamy et al.CVPR 2022 · 10 citations
- HSVLT: Hierarchical Scale-Aware Vision-Language Transformer for Multi-Label Image ClassificationShuyi Ouyang, Hongyi Wang, Ziwei Niu, Zhenjia Bai et al.ACM MM 2023 · 5 citations
- Order-Prompted Tag Sequence Generation for Video TaggingZongyang Ma, Ziqi Zhang, Yuxin Chen, Zhongang Qi et al.ICCV 2023 · 3 citations
- CDDFuse: Correlation-Driven Dual-Branch Feature Decomposition for Multi-Modality Image FusionZixiang Zhao, Haowen Bai, Jiangshe Zhang, Yulun Zhang et al.CVPR 2023
Related papers
- Transformer-based Dual Relation Graph for Multi-label Image RecognitionJiawei Zhao, Ke Yan, Yifan Zhao, Xiaowei Guo et al.ICCV 2021 · 109 citations
- General Multi-Label Image Classification With TransformersJack Lanchantin, Tianlu Wang, Vicente Ordonez, Yanjun QiCVPR 2021
- Two-Stream Transformer for Multi-Label Image ClassificationXuelin Zhu, Jiuxin Cao, Jiawei Ge, Weijia Liu et al.ACM MM 2022 · 37 citations
- Modular Graph Transformer Networks for Multi-Label Image ClassificationHoang D. Nguyen, Xuan-Son Vu, Duc-Trong LeAAAI 2021 · 78 citations
- Improving Intra- and Inter-Modality Visual Relation for Image CaptioningYong Wang, Wenkai Zhang, Qing Liu, Zhengyuan Zhang et al.ACM MM 2020 · 24 citations
