General Multi-Label Image Classification With Transformers
Jack Lanchantin, Tianlu Wang, Vicente Ordonez, Yanjun Qi
摘要
Multi-label image classification is the task of predicting a set of labels corresponding to objects, attributes or other entities present in an image. In this work we propose the Classification Transformer (C-Tran), a general framework for multi-label image classification that leverages Transformers to exploit the complex dependencies among visual features and labels. Our approach consists of a Transformer encoder trained to predict a set of target labels given an input set of masked labels, and visual features from a convolutional neural network. A key ingredient of our method is a label mask training objective that uses a ternary encoding scheme to represent the state of the labels as positive, negative, or unknown during training. Our model shows state-of-the-art performance on challenging datasets such as COCO and Visual Genome. Moreover, because our model explicitly represents the uncertainty of labels during training, it is more general by allowing us to produce improved results for images with partial or extra label annotations during inference. We demonstrate this additional capability in the COCO, Visual Genome, News-500, and CUB image datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper52
- Instance-level Image Retrieval using Reranking TransformersFuwen Tan, Jiangbo Yuan, Vicente OrdonezICCV 2021 · 被引用 116 次
- Transformer-based Dual Relation Graph for Multi-label Image RecognitionJiawei Zhao, Ke Yan, Yifan Zhao, Xiaowei Guo 等ICCV 2021 · 被引用 109 次
- Integrative Few-Shot Learning for Classification and SegmentationDahyun Kang, Minsu ChoCVPR 2022 · 被引用 76 次
- Open-Vocabulary Multi-Label Classification via Multi-Modal Knowledge TransferSunan He, Taian Guo, Tao Dai, Ruizhi Qiao 等AAAI 2023 · 被引用 76 次
- Incomplete Multi-View Multi-Label Learning via Label-Guided Masked View- and Category-Aware TransformersChengliang Liu, Jie Wen, Xiaoling Luo, Yong XuAAAI 2023 · 被引用 68 次
它引用的顶会 Paper3
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Concept Bottleneck ModelsPang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann 等ICML 2020 · 被引用 1,233 次
- Learning Semantic-Specific Graph Representation for Multi-Label Image RecognitionTianshui Chen, Muxin Xu, Xiaolu Hui, Hefeng Wu 等ICCV 2019 · 被引用 347 次
相关 Paper
- Structured Semantic Transfer for Multi-Label Recognition with Partial LabelsTianshui Chen, Tao Pu, Hefeng Wu, Yuan Xie 等AAAI 2022 · 被引用 81 次
- M3TR: Multi-modal Multi-label Recognition with TransformerJiawei Zhao, Yifan Zhao, Jia LiACM MM 2021 · 被引用 45 次
- Modular Graph Transformer Networks for Multi-Label Image ClassificationHoang D. Nguyen, Xuan-Son Vu, Duc-Trong LeAAAI 2021 · 被引用 78 次
- Two-Stream Transformer for Multi-Label Image ClassificationXuelin Zhu, Jiuxin Cao, Jiawei Ge, Weijia Liu 等ACM MM 2022 · 被引用 37 次
- Multimodal Contrastive Training for Visual Representation LearningXin Yuan, Zhe Lin, Jason Kuen, Jianming Zhang 等CVPR 2021
