Learning Clustering-based Prototypes for Compositional Zero-Shot Learning
Hongyu Qu, Jianan Wei, Xiangbo Shu, Wenguan Wang
Abstract
Learning primitive (i.e., attribute and object) concepts from seen compositions is the primary challenge of Compositional Zero-Shot Learning (CZSL). Existing CZSL solutions typically rely on oversimplified data assumptions, e.g., modeling each primitive with a single centroid primitive representation, ignoring the natural diversities of the attribute (resp. object) when coupled with different objects (resp. attribute). In this work, we develop CLUSPRO, a robust clusteringbased prototype mining framework for CZSL that defines the conceptual boundaries of primitives through a set of diversified prototypes. Specifically, CLUSPRO conducts within-primitive clustering on the embedding space for automatically discovering and dynamically updating prototypes. These representative prototypes are subsequently used to repaint a well-structured and independent primitive embedding space, ensuring intra-primitive separation and inter-primitive decorrelation through prototype-based contrastive learning and decorrelation learning. Moreover, CLUSPRO efficiently performs prototype clustering in a nonparametric fashion without the introduction of additional learnable parameters or computational budget during testing. Experiments on three benchmarks demonstrate CLUSPRO outperforms various top-leading CZSL solutions under both closed-world and open-world settings. Our code is available at CLUSPRO. * Equal contribution † Corresponding author 1 Given that CLIP might be exposed to certain unseen compositions during pre-training, we provide detailed data overlap discussion in §G of Appendix.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 783701d1-2201-4a96-9a62-5f4a5f6cee73Cited by top-tier papers14
- Vision-centric Token Compression in Large Language ModelLing Xing, Alex Jinpeng Wang, Rui Yan, Xiangbo Shu et al.NeurIPS 2025 · 32 citations
- OmniGaze: Reward-inspired Generalizable Gaze Estimation in the WildHongyu Qu, Jianan Wei, Xiangbo Shu, Yazhou Yao et al.NeurIPS 2025 · 15 citations
- Multi-Cache Enhanced Prototype Learning for Test-Time Generalization of Vision-Language ModelsXinyu Chen, Haotian Zhai, Can Zhang, Xiupeng Shi et al.ICCV 2025 · 2 citations
- A Conditional Probability Framework for Compositional Zero-Shot LearningPeng Wu, Qiuxia Lai, Hao Fang, Guo-Sen Xie et al.ICCV 2025 · 2 citations
- Incentivizing Generative Zero-Shot Learning via Outcome-Reward Reinforcement Learning with Visual CuesWenjin Hou, Xiaoxiao Sun, Hehe FanCVPR 2026 · 1 citation
Builds on45
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
Related papers
- Compositional Zero-Shot Learning with Contextualized Cues and Adaptive Contrastive TrainingYun Li, Lina Yao, Zhe LiuACM MM 2025
- Leveraging Sub-class Discimination for Compositional Zero-Shot LearningXiaoming Hu, Zilei WangAAAI 2023 · 21 citations
- Independent Prototype Propagation for Zero-Shot CompositionalityFrank Ruis, Gertjan J. Burghouts, Doina BucurNeurIPS 2021 · 77 citations
- LOGICZSL: Exploring Logic-induced Representation for Compositional Zero-shot LearningPeng Wu, Xiankai Lu, Hao Hu, Yongqin Xian et al.CVPR 2025
- Learning Visual Proxy for Compositional Zero-Shot LearningShiyu Zhang, Cheng Yan, Yang Liu, Chenchen Jing et al.ICCV 2025 · 1 citation
