Learning Cross-Aligned Latent Embeddings for Zero-Shot Cross-Modal Retrieval
Kaiyi Lin, Xing Xu, Lianli Gao, Zheng Wang, Heng Tao Shen
摘要
Zero-Shot Cross-Modal Retrieval (ZS-CMR) is an emerging research hotspot that aims to retrieve data of new classes across different modality data. It is challenging for not only the heterogeneous distributions across different modalities, but also the inconsistent semantics across seen and unseen classes. A handful of recently proposed methods typically borrow the idea from zero-shot learning, i.e., exploiting word embeddings of class labels (i.e., class-embeddings) as common semantic space, and using generative adversarial network (GAN) to capture the underlying multimodal data structures, as well as strengthen relations between input data and semantic space to generalize across seen and unseen classes. In this paper, we propose a novel method termed Learning Cross-Aligned Latent Embeddings (LCALE) as an alternative to these GAN based methods for ZS-CMR. Unlike using the class-embeddings as the semantic space, our method seeks for a shared low-dimensional latent space of input multimodal features and class-embeddings by modality-specific variational autoencoders. Notably, we align the distributions learned from multimodal input features and from class-embeddings to construct latent embeddings that contain the essential cross-modal correlation associated with unseen classes. Effective cross-reconstruction and cross-alignment criterions are further developed to preserve class-discriminative information in latent space, which benefits the efficiency for retrieval and enable the knowledge transfer to unseen classes. We evaluate our model using four benchmark datasets on image-text retrieval tasks and one large-scale dataset on image-sketch retrieval tasks. The experimental results show that our method establishes the new state-of-the-art performance for both tasks on all datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Variational Interaction Information Maximization for Cross-domain DisentanglementHyeongJoo Hwang, Geon-Hyeong Kim, Seunghoon Hong, Kee-Eung KimNeurIPS 2020 · 被引用 65 次
- TVT: Three-Way Vision Transformer through Multi-Modal Hypersphere Learning for Zero-Shot Sketch-Based Image RetrievalJialin Tian, Xing Xu, Fumin Shen, Yang Yang 等AAAI 2022 · 被引用 54 次
- Learned Data-aware Image Representations of Line Charts for Similarity SearchYuyu Luo, Yihui Zhou, Nan Tang, Guoliang Li 等SIGMOD 2023 · 被引用 15 次
- Operator SVD with Neural Networks via Nested Low-Rank ApproximationJongha Jon Ryu, Xiangxiang Xu, Hasan Sabri Melihcan Erol, Yuheng Bu 等ICML 2024 · 被引用 10 次
- Seek Commonality but Preserve Differences: Dissected Dynamics Modeling for Multi-modal Visual RLYangru Huang, Peixi Peng, Yifan Zhao, Guangyao Chen 等NeurIPS 2024 · 被引用 3 次
相关 Paper
- Correlated Features Synthesis and Alignment for Zero-shot Cross-modal RetrievalXing Xu, Kaiyi Lin, Huimin Lu, Lianli Gao 等SIGIR 2020 · 被引用 22 次
- Multimodal Disentanglement Variational AutoEncoders for Zero-Shot Cross-Modal RetrievalJialin Tian, Kai Wang, Xing Xu, Zuo Cao 等SIGIR 2022 · 被引用 19 次
- Learning Aligned Cross-Modal Representation for Generalized Zero-Shot ClassificationZhiyu Fang, Xiaobin Zhu, Chun Yang, Zheng Han 等AAAI 2022 · 被引用 26 次
- A Variational Autoencoder with Deep Embedding Model for Generalized Zero-Shot LearningPeirong Ma, Xiao HuAAAI 2020 · 被引用 43 次
- Exploring Graph-Structured Semantics for Cross-Modal RetrievalLei Zhang, Leiting Chen, Chuan Zhou, Fan Yang 等ACM MM 2021 · 被引用 14 次
