Learning Aligned Cross-Modal Representation for Generalized Zero-Shot Classification
Zhiyu Fang, Xiaobin Zhu, Chun Yang, Zheng Han, Jingyan Qin, Xu-Cheng Yin
摘要
Learning a common latent embedding by aligning the latent spaces of cross-modal autoencoders is an effective strategy for Generalized Zero-Shot Classification (GZSC). However, due to the lack of fine-grained instance-wise annotations, it still easily suffer from the domain shift problem for the discrepancy between the visual representation of diversified images and the semantic representation of fixed attributes. In this paper, we propose an innovative autoencoder network by learning Aligned Cross-Modal Representations (dubbed ACMR) for GZSC. Specifically, we propose a novel Vision-Semantic Alignment (VSA) method to strengthen the alignment of cross-modal latent features on the latent subspaces guided by a learned classifier. In addition, we propose a novel Information Enhancement Module (IEM) to reduce the possibility of latent variables collapse meanwhile encouraging the discriminative ability of latent variables. Extensive experiments on publicly available datasets demonstrate the state-of-the-art performance of our method.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Bayesian Cross-Modal Alignment Learning for Few-Shot Out-of-Distribution GeneralizationLin Zhu, Xinbing Wang, Chenghu Zhou, Nanyang YeAAAI 2023 · 被引用 8 次
- Amplifying Prominent Representations in Multimodal Learning via Variational Dirichlet ProcessTsai Hor Chan, Feng Wu, Yihang Chen, Guosheng Yin 等NeurIPS 2025
它引用的顶会 Paper13
- Semi-Supervised Domain Adaptation via Minimax EntropyKuniaki Saito, Donghyun Kim, Stan Sclaroff, Trevor Darrell 等ICCV 2019 · 被引用 725 次
- Attribute Prototype Network for Zero-Shot LearningWenjia Xu, Yongqin Xian, Jiuniu Wang, Bernt Schiele 等NeurIPS 2020 · 被引用 392 次
- Adaptive Boundary Proposal Network for Arbitrary Shape Text DetectionShi-Xue Zhang, Xiaobin Zhu, Chun Yang, Hongfa Wang 等ICCV 2021 · 被引用 112 次
- Generalized Zero-Shot Learning via Disentangled RepresentationXiangyu Li, Zhe Xu, Kun Wei, Cheng DengAAAI 2021 · 被引用 88 次
- A Variational Autoencoder with Deep Embedding Model for Generalized Zero-Shot LearningPeirong Ma, Xiao HuAAAI 2020 · 被引用 43 次
相关 Paper
- Learning Cross-Aligned Latent Embeddings for Zero-Shot Cross-Modal RetrievalKaiyi Lin, Xing Xu, Lianli Gao, Zheng Wang 等AAAI 2020 · 被引用 50 次
- Distinguishing Unseen from Seen for Generalized Zero-shot LearningHongzu Su, Jingjing Li, Zhi Chen, Lei Zhu 等CVPR 2022 · 被引用 40 次
- Learning Modality-Invariant Latent Representations for Generalized Zero-shot LearningJingjing Li, Mengmeng Jing, Lei Zhu, Zhengming Ding 等ACM MM 2020 · 被引用 35 次
- Task-Independent Knowledge Makes for Transferable Representations for Generalized Zero-Shot LearningChaoqun Wang, Xuejin Chen, Shaobo Min, Xiaoyan Sun 等AAAI 2021 · 被引用 22 次
- Multimodal Disentanglement Variational AutoEncoders for Zero-Shot Cross-Modal RetrievalJialin Tian, Kai Wang, Xing Xu, Zuo Cao 等SIGIR 2022 · 被引用 19 次
