Bridging the Modality Gap in Compositional Zero-Shot Learning via Sparse Alignment and Unimodal Memory Bank
Yang Zhang, Zhixiang Chi, Xudong Yan, Yang Wang, Songhe Feng
Abstract
Compositional Zero-Shot Learning (CZSL) aims to recognize unseen attribute-object compositions with learned primitives (attribute and object) knowledge from seen compositions. While previous approaches gain their notable performance through the powerful cross-modal alignment of CLIP, they often overlook the modality gap, an inherent constraint stemming from information-imbalanced training data. In this work, we propose SAM, a novel parse lignment and Unimoal emory Bank to effectively bridging modality gap for CZSL. Specifically, we conduct that links textual representations directly to their semantically pertinent visual patches. This direct linking serves to prune redundant visual data and counter the information imbalance in image-text pairs. Subsequently, with the sparsely aligned visual information as its guidance, the module adaptively fuses these critical cues into a unified representation. Finally, we introduce a that stores samples from both seen and unseen compositions. This bank serves a dual purpose: it bypasses the modality gap through visual-only classification and concurrently strengthens generalization to unseen compositions. Experiments on three benchmarks demonstrate that our method gains significant improvements over CLIP-based methods under closed-world and open-world settings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on35
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation LearningWeixin Liang, Yuhui Zhang, Yongchan Kwon, Serena Yeung et al.NeurIPS 2022 · 834 citations
- SuS-X: Training-Free Name-Only Transfer of Vision-Language ModelsVishaal Udandarao, Ankush Gupta, Samuel AlbanieICCV 2023 · 160 citations
Related papers
- Learning Visual Proxy for Compositional Zero-Shot LearningShiyu Zhang, Cheng Yan, Yang Liu, Chenchen Jing et al.ICCV 2025 · 1 citation
- TOMCAT: Test-time Comprehensive Knowledge Accumulation for Compositional Zero-Shot LearningXudong Yan, Songhe FengNeurIPS 2025
- Compositional Zero-Shot Learning with Contextualized Cues and Adaptive Contrastive TrainingYun Li, Lina Yao, Zhe LiuACM MM 2025
- LOGICZSL: Exploring Logic-induced Representation for Compositional Zero-shot LearningPeng Wu, Xiankai Lu, Hao Hu, Yongqin Xian et al.CVPR 2025
- ProCC: Progressive Cross-Primitive Compatibility for Open-World Compositional Zero-Shot LearningFushuo Huo, Wenchao Xu, Song Guo, Jingcai Guo et al.AAAI 2024 · 17 citations
