Learning Cross-Aligned Latent Embeddings for Zero-Shot Cross-Modal Retrieval
Kaiyi Lin, Xing Xu, Lianli Gao, Zheng Wang, Heng Tao Shen
Abstract
Zero-Shot Cross-Modal Retrieval (ZS-CMR) is an emerging research hotspot that aims to retrieve data of new classes across different modality data. It is challenging for not only the heterogeneous distributions across different modalities, but also the inconsistent semantics across seen and unseen classes. A handful of recently proposed methods typically borrow the idea from zero-shot learning, i.e., exploiting word embeddings of class labels (i.e., class-embeddings) as common semantic space, and using generative adversarial network (GAN) to capture the underlying multimodal data structures, as well as strengthen relations between input data and semantic space to generalize across seen and unseen classes. In this paper, we propose a novel method termed Learning Cross-Aligned Latent Embeddings (LCALE) as an alternative to these GAN based methods for ZS-CMR. Unlike using the class-embeddings as the semantic space, our method seeks for a shared low-dimensional latent space of input multimodal features and class-embeddings by modality-specific variational autoencoders. Notably, we align the distributions learned from multimodal input features and from class-embeddings to construct latent embeddings that contain the essential cross-modal correlation associated with unseen classes. Effective cross-reconstruction and cross-alignment criterions are further developed to preserve class-discriminative information in latent space, which benefits the efficiency for retrieval and enable the knowledge transfer to unseen classes. We evaluate our model using four benchmark datasets on image-text retrieval tasks and one large-scale dataset on image-sketch retrieval tasks. The experimental results show that our method establishes the new state-of-the-art performance for both tasks on all datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e5a82dd7-4911-4655-a3f9-baec13cd1f40Cited by top-tier papers6
- Variational Interaction Information Maximization for Cross-domain DisentanglementHyeongJoo Hwang, Geon-Hyeong Kim, Seunghoon Hong, Kee-Eung KimNeurIPS 2020 · 65 citations
- TVT: Three-Way Vision Transformer through Multi-Modal Hypersphere Learning for Zero-Shot Sketch-Based Image RetrievalJialin Tian, Xing Xu, Fumin Shen, Yang Yang et al.AAAI 2022 · 54 citations
- Learned Data-aware Image Representations of Line Charts for Similarity SearchYuyu Luo, Yihui Zhou, Nan Tang, Guoliang Li et al.SIGMOD 2023 · 15 citations
- Operator SVD with Neural Networks via Nested Low-Rank ApproximationJongha Jon Ryu, Xiangxiang Xu, Hasan Sabri Melihcan Erol, Yuheng Bu et al.ICML 2024 · 10 citations
- Seek Commonality but Preserve Differences: Dissected Dynamics Modeling for Multi-modal Visual RLYangru Huang, Peixi Peng, Yifan Zhao, Guangyao Chen et al.NeurIPS 2024 · 3 citations
Related papers
- Correlated Features Synthesis and Alignment for Zero-shot Cross-modal RetrievalXing Xu, Kaiyi Lin, Huimin Lu, Lianli Gao et al.SIGIR 2020 · 22 citations
- Multimodal Disentanglement Variational AutoEncoders for Zero-Shot Cross-Modal RetrievalJialin Tian, Kai Wang, Xing Xu, Zuo Cao et al.SIGIR 2022 · 19 citations
- Learning Aligned Cross-Modal Representation for Generalized Zero-Shot ClassificationZhiyu Fang, Xiaobin Zhu, Chun Yang, Zheng Han et al.AAAI 2022 · 26 citations
- A Variational Autoencoder with Deep Embedding Model for Generalized Zero-Shot LearningPeirong Ma, Xiao HuAAAI 2020 · 43 citations
- Exploring Graph-Structured Semantics for Cross-Modal RetrievalLei Zhang, Leiting Chen, Chuan Zhou, Fan Yang et al.ACM MM 2021 · 14 citations
