Discriminative Region-based Multi-Label Zero-Shot Learning
Sanath Narayan, Akshita Gupta, Salman H. Khan, Fahad Shahbaz Khan, Ling Shao, Mubarak Shah
摘要
Multi-label zero-shot learning (ZSL) is a more realistic counter-part of standard single-label ZSL since several objects can co-exist in a natural image. However, the occurrence of multiple objects complicates the reasoning and requires region-specific processing of visual features to preserve their contextual cues. We note that the best existing multi-label ZSL method takes a shared approach towards attending to region features with a common set of attention maps for all the classes. Such shared maps lead to diffused attention, which does not discriminatively focus on relevant locations when the number of classes are large. Moreover, mapping spatially-pooled visual features to the class semantics leads to inter-class feature entanglement, thus hampering the classification. Here, we propose an alternate approach towards region-based discriminability-preserving multi-label zero-shot classification. Our approach maintains the spatial resolution to preserve region-level characteristics and utilizes a bi-level attention module (BiAM) to enrich the features by incorporating both region and scene context information. The enriched region-level features are then mapped to the class semantics and only their class predictions are spatially pooled to obtain image-level predictions, thereby keeping the multi-class features disentangled. Our approach sets a new state of the art on two large-scale multi-label zero-shot benchmarks: NUS-WIDE and Open Images. On NUS-WIDE, our approach achieves an absolute gain of 6.9% mAP for ZSL, compared to the best published results. Source code is available at https://github.com/akshitac8/BiAM.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- DualCoOp: Fast Adaptation to Multi-Label Recognition with Limited AnnotationsXimeng Sun, Ping Hu, Kate SaenkoNeurIPS 2022 · 被引用 199 次
- Open-Vocabulary Multi-Label Classification via Multi-Modal Knowledge TransferSunan He, Taian Guo, Tao Dai, Ruizhi Qiao 等AAAI 2023 · 被引用 76 次
- TagCLIP: A Local-to-Global Framework to Enhance Open-Vocabulary Multi-Label Classification of CLIP without TrainingYuqi Lin, Minghao Chen, Kaipeng Zhang, Hengjia Li 等AAAI 2024 · 被引用 39 次
- Simple Image-Level Classification Improves Open-Vocabulary Object DetectionRuohuan Fang, Guansong Pang, Xiao BaiAAAI 2024 · 被引用 26 次
- Language-Driven Cross-Modal Classifier for Zero-Shot Multi-Label Image RecognitionYicheng Liu, Jie Wen, Chengliang Liu, Xiaozhao Fang 等ICML 2024 · 被引用 7 次
它引用的顶会 Paper7
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- CCNet: Criss-Cross Attention for Semantic SegmentationZilong Huang, Xinggang Wang, Lichao Huang, Chang Huang 等ICCV 2019 · 被引用 2,972 次
- Cross-Modality Attention with Semantic Graph Embedding for Multi-Label ClassificationRenchun You, Zhiyao Guo, Lei Cui, Xiang Long 等AAAI 2020 · 被引用 221 次
- Semantic Diversity Learning for Zero-Shot Multi-label ClassificationAvi Ben-Cohen, Nadav Zamir, Emanuel Ben Baruch, Itamar Friedman 等ICCV 2021 · 被引用 47 次
- A Shared Multi-Attention Framework for Multi-Label Zero-Shot LearningDat Huynh, Ehsan ElhamifarCVPR 2020
相关 Paper
- Epsilon: Exploring Comprehensive Visual-Semantic Projection for Multi-Label Zero-Shot LearningZiming Liu, Jingcai Guo, Song Guo, Xiaocheng LuAAAI 2025 · 被引用 6 次
- Goal-Oriented Gaze Estimation for Zero-Shot LearningYang Liu, Lei Zhou, Xiao Bai, Yifei Huang 等CVPR 2021
- Attribute Attention for Semantic Disambiguation in Zero-Shot LearningYang Liu, Jishun Guo, Deng Cai, Xiaofei HeICCV 2019 · 被引用 163 次
- Generalized Zero-Shot Learning via Disentangled RepresentationXiangyu Li, Zhe Xu, Kun Wei, Cheng DengAAAI 2021 · 被引用 88 次
- Leveraging Sub-class Discimination for Compositional Zero-Shot LearningXiaoming Hu, Zilei WangAAAI 2023 · 被引用 21 次
