Enhancing Domain-Invariant Parts for Generalized Zero-Shot Learning
Yang Zhang, Songhe Feng
Abstract
Generalized Zero-Shot Learning (GZSL) aims to recognize unseen classes, which are not observable during training, with auxiliary semantic information, e.g., attributes. Attributes play an important role in GZSL due to their ability to provide rich semantic information and visual guidance. In this paper, we will review the characteristics of attributes. First, the visual appearance of the same attribute varies greatly between different classes, which leads to the projection domain shift problem and inaccurate attribute localization. Second, the predefined semantic prototypes are not able to faithfully represent the credible visual information of each sample, which leads to suboptimal visual-semantic alignment. Therefore, we propose a novel framework called Enhancing doMain-invariant Parts (EMP) to solve above issues. To be specific, we mitigate the projection domain shift problem by the feature disentanglement technology in domain generalization, which can disentangle the attribute-based visual features into domain-invariant and domain-specific parts. So that the model can pay more attention to the essential parts of attributes rather than using all the information learned from the seen classes to identify the unseen classes. Then we achieve better visual-semantic alignment by refining the predefined semantic prototypes with the extracted credible visual information from corresponding sample. To make the extracted visual information be more in line with the image content rather than overfitting the semantic prototype, we draw on the idea of self-paced learning to help the model learn attributes from easy to complex. Experimental results show that our method achieves a new state-of-the-art performance on three generalized zero-shot learning benchmarks.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers3
- Visual-Semantic Decomposition and Partial Alignment for Document-based Zero-Shot LearningXiangyan Qu, Jing Yu, Keke Gai, Jiamin Zhuang et al.ACM MM 2024 · 5 citations
- Bridging the Modality Gap in Compositional Zero-Shot Learning via Sparse Alignment and Unimodal Memory BankYang Zhang, Zhixiang Chi, Xudong Yan, Yang Wang et al.CVPR 2026
- TOMCAT: Test-time Comprehensive Knowledge Accumulation for Compositional Zero-Shot LearningXudong Yan, Songhe FengNeurIPS 2025
Related papers
- Dual Progressive Prototype Network for Generalized Zero-Shot LearningChaoqun Wang, Shaobo Min, Xuejin Chen, Xiaoyan Sun et al.NeurIPS 2021 · 72 citations
- Semantics Disentangling for Generalized Zero-Shot LearningZhi Chen, Yadan Luo, Ruihong Qiu, Sen Wang et al.ICCV 2021 · 143 citations
- A Variational Autoencoder with Deep Embedding Model for Generalized Zero-Shot LearningPeirong Ma, Xiao HuAAAI 2020 · 43 citations
- SAGE: Structured Attribute-Guided Enhancement for GZSLZao Zhang, Liguo Sun, Pin LyuAAAI 2026
- Semantic-guided Reinforced Region Embedding for Generalized Zero-Shot LearningJiannan Ge, Hongtao Xie, Shaobo Min, Yongdong ZhangAAAI 2021 · 36 citations
