Orderless Recurrent Models for Multi-Label Classification
Vacit Oguz Yazici, Abel Gonzalez-Garcia, Arnau Ramisa, Bartlomiej Twardowski, Joost van de Weijer
Abstract
Recurrent neural networks (RNN) are popular for many computer vision tasks, including multi-label classification. Since RNNs produce sequential outputs, labels need to be ordered for the multi-label classification task. Current approaches sort labels according to their frequency, typically ordering them in either rare-first or frequent-first. These imposed orderings do not take into account that the natural order to generate the labels can change for each image, e.g. first the dominant object before summing up the smaller objects in the image. Therefore, we propose ways to dynamically order the ground truth labels with the predicted label sequence. This allows for faster training of more optimal LSTM models. Analysis evidences that our method does not suffer from duplicate generation, something which is common for other models. Furthermore, it outperforms other CNN-RNN models, and we show that a standard architecture of an image encoder and language decoder trained with our proposed loss obtains the state-of-the-art results on the challenging MS-COCO, WIDER Attribute and PA-100K and competitive results on NUS-WIDE.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c56d3925-79ec-453d-89e9-9956d6cbc5c7Cited by top-tier papers28
- DualCoOp: Fast Adaptation to Multi-Label Recognition with Limited AnnotationsXimeng Sun, Ping Hu, Kate SaenkoNeurIPS 2022 · 199 citations
- Transformer-based Dual Relation Graph for Multi-label Image RecognitionJiawei Zhao, Ke Yan, Yifan Zhao, Xiaowei Guo et al.ICCV 2021 · 109 citations
- Large Loss Matters in Weakly Supervised Multi-Label ClassificationYoungwook Kim, Jae-Myung Kim, Zeynep Akata, Jungwoo LeeCVPR 2022 · 68 citations
- Discriminative Region-based Multi-Label Zero-Shot LearningSanath Narayan, Akshita Gupta, Salman H. Khan, Fahad Shahbaz Khan et al.ICCV 2021 · 62 citations
- NeuralTailor: reconstructing sewing pattern structures from 3D point clouds of garmentsMaria Korosteleva, Sung-Hee LeeSIGGRAPH 2022 · 45 citations
Related papers
- Expressing Objects Just Like Words: Recurrent Visual Embedding for Image-Text MatchingTianlang Chen, Jiebo LuoAAAI 2020 · 71 citations
- Order-Free Learning Alleviating Exposure Bias in Multi-Label ClassificationChe-Ping Tsai, Hung-yi LeeAAAI 2020 · 36 citations
- RATT: Recurrent Attention to Transient Tasks for Continual Image CaptioningRiccardo Del Chiaro, Bartlomiej Twardowski, Andrew D. Bagdanov, Joost van de WeijerNeurIPS 2020 · 55 citations
- Rank & Sort Loss for Object Detection and Instance SegmentationKemal Oksuz, Baris Can Cam, Emre Akbas, Sinan KalkanICCV 2021 · 49 citations
- Exploring Overall Contextual Information for Image Captioning in Human-Like Cognitive StyleHongwei Ge, Zehang Yan, Kai Zhang, Mingde Zhao et al.ICCV 2019 · 25 citations
