Bridging the Modality Gap in Compositional Zero-Shot Learning via Sparse Alignment and Unimodal Memory Bank
Yang Zhang, Zhixiang Chi, Xudong Yan, Yang Wang, Songhe Feng
摘要
Compositional Zero-Shot Learning (CZSL) aims to recognize unseen attribute-object compositions with learned primitives (attribute and object) knowledge from seen compositions. While previous approaches gain their notable performance through the powerful cross-modal alignment of CLIP, they often overlook the modality gap, an inherent constraint stemming from information-imbalanced training data. In this work, we propose SAM, a novel parse lignment and Unimoal emory Bank to effectively bridging modality gap for CZSL. Specifically, we conduct that links textual representations directly to their semantically pertinent visual patches. This direct linking serves to prune redundant visual data and counter the information imbalance in image-text pairs. Subsequently, with the sparsely aligned visual information as its guidance, the module adaptively fuses these critical cues into a unified representation. Finally, we introduce a that stores samples from both seen and unseen compositions. This bank serves a dual purpose: it bypasses the modality gap through visual-only classification and concurrently strengthens generalization to unseen compositions. Experiments on three benchmarks demonstrate that our method gains significant improvements over CLIP-based methods under closed-world and open-world settings.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper35
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation LearningWeixin Liang, Yuhui Zhang, Yongchan Kwon, Serena Yeung 等NeurIPS 2022 · 被引用 834 次
- SuS-X: Training-Free Name-Only Transfer of Vision-Language ModelsVishaal Udandarao, Ankush Gupta, Samuel AlbanieICCV 2023 · 被引用 160 次
相关 Paper
- Learning Visual Proxy for Compositional Zero-Shot LearningShiyu Zhang, Cheng Yan, Yang Liu, Chenchen Jing 等ICCV 2025 · 被引用 1 次
- TOMCAT: Test-time Comprehensive Knowledge Accumulation for Compositional Zero-Shot LearningXudong Yan, Songhe FengNeurIPS 2025
- Compositional Zero-Shot Learning with Contextualized Cues and Adaptive Contrastive TrainingYun Li, Lina Yao, Zhe LiuACM MM 2025
- LOGICZSL: Exploring Logic-induced Representation for Compositional Zero-shot LearningPeng Wu, Xiankai Lu, Hao Hu, Yongqin Xian 等CVPR 2025
- ProCC: Progressive Cross-Primitive Compatibility for Open-World Compositional Zero-Shot LearningFushuo Huo, Wenchao Xu, Song Guo, Jingcai Guo 等AAAI 2024 · 被引用 17 次
