Learning Visual Proxy for Compositional Zero-Shot Learning
Shiyu Zhang, Cheng Yan, Yang Liu, Chenchen Jing, Lei Zhou, Wenjun Wang
Abstract
Compositional Zero-Shot Learning (CZSL) aims to recognize novel attribute-object compositions by leveraging knowledge from seen compositions. Current methods align textual prototypes with visual features via Vision-Language Models (VLMs), but suffer from two limitations: (1) modality gaps hinder the discrimination of semantically similar pairs, and (2) single-modal textual prototypes lack finegrained visual cues. In this paper, we introduce Visual Proxy Learning, a method that reduces modality gaps and enhances compositional generalization. We initialize visual proxies for attributes, objects, and their compositions using text representations and optimize the visual space to capture fine-grained cues, improving visual representations. Additionally, we propose Cross-Modal Joint Learning (CMJL), which imposes cross-modal constraints between the textimage and fine-grained visual spaces, improving generalization for unseen compositions and discriminating similar pairs. Experiments show state-of-the-art performance in closed-world scenarios and competitive results in openworld settings across four CZSL benchmarks, demonstrating the effectiveness of our approach in compositional generalization. The code will be available at https:// github.com/codefish12-09/VP_CMJL.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f643e6e1-1a07-4995-867c-ed2085e494b6Cited by top-tier papers1
Ask how each one uses itBuilds on23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image ClassificationChun-Fu (Richard) Chen, Quanfu Fan, Rameswar PandaICCV 2021 · 2,072 citations
- AdaptFormer: Adapting Vision Transformers for Scalable Visual RecognitionShoufa Chen, Chongjian Ge, Zhan Tong, Jiangliu Wang et al.NeurIPS 2022 · 1,291 citations
- Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation LearningWeixin Liang, Yuhui Zhang, Yongchan Kwon, Serena Yeung et al.NeurIPS 2022 · 834 citations
- Task-Driven Modular Networks for Zero-Shot Compositional LearningSenthil Purushwalkam, Maximilian Nickel, Abhinav Gupta, Marc'Aurelio RanzatoICCV 2019 · 222 citations
Related papers
- TOMCAT: Test-time Comprehensive Knowledge Accumulation for Compositional Zero-Shot LearningXudong Yan, Songhe FengNeurIPS 2025
- Troika: Multi-Path Cross-Modal Traction for Compositional Zero-Shot LearningSiteng Huang, Biao Gong, Yutong Feng, Min Zhang et al.CVPR 2024
- Decomposed Soft Prompt Guided Fusion Enhancing for Compositional Zero-Shot LearningXiaocheng Lu, Song Guo, Ziming Liu, Jingcai GuoCVPR 2023
- Bridging the Modality Gap in Compositional Zero-Shot Learning via Sparse Alignment and Unimodal Memory BankYang Zhang, Zhixiang Chi, Xudong Yan, Yang Wang et al.CVPR 2026
- Learning Conditional Attributes for Compositional Zero-Shot LearningQingsheng Wang, Lingqiao Liu, Chenchen Jing, Hao Chen et al.CVPR 2023
