Semantic-guided Reinforced Region Embedding for Generalized Zero-Shot Learning
Jiannan Ge, Hongtao Xie, Shaobo Min, Yongdong Zhang
Abstract
Generalized zero-shot Learning (GZSL) aims to recognize images from either seen or unseen domain, mainly by learning a joint embedding space to associate image features with the corresponding category descriptions. Recent methods have proved that localizing important object regions can effectively bridge the semantic-visual gap. However, these are all based on one-off visual localizers, lacking of interpretability and flexibility. In this paper, we propose a novel Semantic-guided Reinforced Region Embedding (SR2E) network that can localize important objects in the long-term interests to construct semantic-visual embedding space. SR2E consists of Reinforced Region Module (R2M) and Semantic Alignment Module (SAM). First, without the annotated bounding box as supervision, R2M encodes the semantic category guidance into the reward and punishment criteria to teach the localizer serialized region searching. Besides, R2M explores different action spaces during the serialized searching path to avoid local optimal localization, which thereby generates discriminative visual features with less redundancy. Second, SAM preserves the semantic relationship into visual features via semantic-visual alignment and designs a domain detector to alleviate the domain confusion. Experiments on four public benchmarks demonstrate that the proposed SR2E is an effective GZSL method with reinforced embedding space, which obtains averaged 6.1% improvements.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- From Two to One: A New Scene Text Recognizer with Visual Language Modeling NetworkYuxin Wang, Hongtao Xie, Shancheng Fang, Jing Wang et al.ICCV 2021 · 184 citations
- Towards Balanced Alignment: Modal-Enhanced Semantic Modeling for Video Moment RetrievalZhihang Liu, Jun Li, Hongtao Xie, Pandeng Li et al.AAAI 2024 · 49 citations
- Neighborhood-Adaptive Structure Augmented Metric LearningPandeng Li, Yan Li, Hongtao Xie, Lei ZhangAAAI 2022 · 29 citations
- Learning Aligned Cross-Modal Representation for Generalized Zero-Shot ClassificationZhiyu Fang, Xiaobin Zhu, Chun Yang, Zheng Han et al.AAAI 2022 · 26 citations
- Distilled Reverse Attention Network for Open-world Compositional Zero-Shot LearningYun Li, Zhe Liu, Saurav Jha, Lina YaoICCV 2023 · 23 citations
Builds on6
- AU-assisted Graph Attention Convolutional Network for Micro-Expression RecognitionHong-Xia Xie, Ling Lo, Hong-Han Shuai, Wen-Huang ChengACM MM 2020 · 189 citations
- Filtration and Distillation: Enhancing Region Attention for Fine-Grained Visual CategorizationChuanbin Liu, Hongtao Xie, Zheng-Jun Zha, Lingfeng Ma et al.AAAI 2020 · 179 citations
- S2SiamFC: Self-supervised Fully Convolutional Siamese Network for Visual TrackingChon-Hou Sio, Yu-Jen Ma, Hong-Han Shuai, Jun-Cheng Chen et al.ACM MM 2020 · 45 citations
- Fine-Grained Generalized Zero-Shot Learning via Dense Attribute-Based AttentionDat Huynh, Ehsan ElhamifarCVPR 2020
- Episode-Based Prototype Generating Network for Zero-Shot LearningYunlong Yu, Zhong Ji, Jungong Han, Zhongfei ZhangCVPR 2020
Related papers
- Self-Supervised Domain-Aware Generative Network for Generalized Zero-Shot LearningJiamin Wu, Tianzhu Zhang, Zheng-Jun Zha, Jiebo Luo et al.CVPR 2020
- A Variational Autoencoder with Deep Embedding Model for Generalized Zero-Shot LearningPeirong Ma, Xiao HuAAAI 2020 · 43 citations
- Task-Independent Knowledge Makes for Transferable Representations for Generalized Zero-Shot LearningChaoqun Wang, Xuejin Chen, Shaobo Min, Xiaoyan Sun et al.AAAI 2021 · 22 citations
- Dual Progressive Prototype Network for Generalized Zero-Shot LearningChaoqun Wang, Shaobo Min, Xuejin Chen, Xiaoyan Sun et al.NeurIPS 2021 · 72 citations
- Adaptive and Generative Zero-Shot LearningYu-Ying Chou, Hsuan-Tien Lin, Tyng-Luh LiuICLR 2021 · 25 citations
