Transformer-based Dual Relation Graph for Multi-label Image Recognition
Jiawei Zhao, Ke Yan, Yifan Zhao, Xiaowei Guo, Feiyue Huang, Jia Li
Abstract
The simultaneous recognition of multiple objects in one image remains a challenging task, spanning multiple events in the recognition field such as various object scales, inconsistent appearances, and confused inter-class relationships. Recent research efforts mainly resort to the statistic label co-occurrences and linguistic word embedding to enhance the unclear semantics. Different from these researches, in this paper, we propose a novel Transformer-based Dual Relation learning framework, constructing complementary relationships by exploring two aspects of correlation, i.e., structural relation graph and semantic relation graph. The structural relation graph aims to capture long-range correlations from object context, by developing a cross-scale transformer-based architecture. The semantic graph dynamically models the semantic meanings of image objects with explicit semantic-aware constraints. In addition, we also incorporate the learnt structural relationship into the semantic graph, constructing a joint relation graph for robust representations. With the collaborative learning of these two effective relation graphs, our approach achieves new state-of-the-art on two popular multi-label recognition benchmarks, i.e. MS-COCO and VOC 2007 dataset.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4f2ac6ad-ed58-409a-b53e-86bdbd20ebffCited by top-tier papers14
- The Surprising Effectiveness of Diffusion Models for Optical Flow and Monocular Depth EstimationSaurabh Saxena, Charles Herrmann, Junhwa Hur, Abhishek Kar et al.NeurIPS 2023 · 160 citations
- SQLdepth: Generalizable Self-Supervised Fine-Structured Monocular Depth EstimationYouhong Wang, Yunji Liang, Hao Xu, Shaohui Jiao et al.AAAI 2024 · 60 citations
- Contextual Debiasing for Visual Recognition with Causal MechanismsRuyang Liu, Hao Liu, Ge Li, Haodi Hou et al.CVPR 2022 · 42 citations
- Self-Supervised Heterogeneous Graph Learning: a Homophily and Heterogeneity ViewYujie Mo, Feiping Nie, Ping Hu, Heng Tao Shen et al.ICLR 2024 · 17 citations
- Holistic Label Correction for Noisy Multi-Label ClassificationXiaobo Xia, Jiankang Deng, Wei Bao, Yuxuan Du et al.ICCV 2023 · 13 citations
Builds on10
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- Learning Semantic-Specific Graph Representation for Multi-Label Image RecognitionTianshui Chen, Muxin Xu, Xiaolu Hui, Hefeng Wu et al.ICCV 2019 · 347 citations
- Cross-Modality Attention with Semantic Graph Embedding for Multi-Label ClassificationRenchun You, Zhiyao Guo, Lei Cui, Xiang Long et al.AAAI 2020 · 221 citations
- Multi-Label Classification with Label Graph SuperimposingYa Wang, Dongliang He, Fu Li, Xiang Long et al.AAAI 2020 · 192 citations
Related papers
- M3TR: Multi-modal Multi-label Recognition with TransformerJiawei Zhao, Yifan Zhao, Jia LiACM MM 2021 · 45 citations
- Dual Graph Convolutional Networks with Transformer and Curriculum Learning for Image CaptioningXinzhi Dong, Chengjiang Long, Wenju Xu, Chunxia XiaoACM MM 2021 · 75 citations
- Two-Stream Transformer for Multi-Label Image ClassificationXuelin Zhu, Jiuxin Cao, Jiawei Ge, Weijia Liu et al.ACM MM 2022 · 37 citations
- Dual-level Collaborative Transformer for Image CaptioningYunpeng Luo, Jiayi Ji, Xiaoshuai Sun, Liujuan Cao et al.AAAI 2021 · 349 citations
- General Multi-Label Image Classification With TransformersJack Lanchantin, Tianlu Wang, Vicente Ordonez, Yanjun QiCVPR 2021
