Learning Modality-Invariant Latent Representations for Generalized Zero-shot Learning
Jingjing Li, Mengmeng Jing, Lei Zhu, Zhengming Ding, Ke Lu, Yang Yang
Abstract
Recently, feature generating methods have been successfully applied to zero-shot learning (ZSL). However, most previous approaches only generate visual representations for zero-shot recognition. In fact, typical ZSL is a classic multi-modal learning protocol which consists of a visual space and a semantic space. In this paper, therefore, we present a new method which can simultaneously generate both visual representations and semantic representations so that the essential multi-modal information associated with unseen classes can be captured. Specifically, we address the most challenging issue in such a paradigm, i.e., how to handle the domain shift and thus guarantee that the learned representations are modality-invariant. To this end, we propose two strategies: 1) leveraging the mutual information between the latent visual representations and the semantic representations; 2) maximizing the entropy of the joint distribution of the two latent representations. By leveraging the two strategies, we argue that the two modalities can be well aligned. At last, extensive experiments on five widely used datasets verify that the proposed method is able to significantly outperform previous the state-of-the-arts.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 2e648246-2d5f-43fa-bd65-0f4f9ca127acCited by top-tier papers5
- Distinguishing Unseen from Seen for Generalized Zero-shot LearningHongzu Su, Jingjing Li, Zhi Chen, Lei Zhu et al.CVPR 2022 · 40 citations
- Mitigating Generation Shifts for Generalized Zero-Shot LearningZhi Chen, Yadan Luo, Sen Wang, Ruihong Qiu et al.ACM MM 2021 · 30 citations
- Learning Aligned Cross-Modal Representation for Generalized Zero-Shot ClassificationZhiyu Fang, Xiaobin Zhu, Chun Yang, Zheng Han et al.AAAI 2022 · 26 citations
- RMIB: Representation Matching Information Bottleneck for Matching Text RepresentationsHaihui Pan, Zhifang Liao, Wenrui Xie, Kun HanICML 2024 · 1 citation
- (ML)2P-Encoder: On Exploration of Channel-Class Correlation for Multi-Label Zero-Shot LearningZiming Liu, Song Guo, Xiaocheng Lu, Jingcai Guo et al.CVPR 2023
Related papers
- Episode-Based Prototype Generating Network for Zero-Shot LearningYunlong Yu, Zhong Ji, Jungong Han, Zhongfei ZhangCVPR 2020
- HSVA: Hierarchical Semantic-Visual Adaptation for Zero-Shot LearningShiming Chen, Guo-Sen Xie, Yang Liu, Qinmu Peng et al.NeurIPS 2021 · 190 citations
- A Variational Autoencoder with Deep Embedding Model for Generalized Zero-Shot LearningPeirong Ma, Xiao HuAAAI 2020 · 43 citations
- Generalized Zero-shot Learning with Multi-source Semantic Embeddings for Scene RecognitionXinhang Song, Haitao Zeng, Sixian Zhang, Luis Herranz et al.ACM MM 2020 · 9 citations
- Evolving Semantic Prototype Improves Generative Zero-Shot LearningShiming Chen, Wenjin Hou, Ziming Hong, Xiaohan Ding et al.ICML 2023 · 33 citations
