Learning Attention as Disentangler for Compositional Zero-Shot Learning
Shaozhe Hao, Kai Han, Kwan-Yee K. Wong
摘要
Compositional zero-shot learning (CZSL) aims at learning visual concepts (i.e., attributes and objects) from seen compositions and combining concept knowledge into unseen compositions. The key to CZSL is learning the disentanglement of the attribute-object composition. To this end, we propose to exploit cross-attentions as compositional disentanglers to learn disentangled concept embeddings. For example, if we want to recognize an unseen composition “yellow flower”, we can learn the attribute concept “yellow” and object concept “flower” from different yellow objects and different flowers respectively. To further constrain the disentanglers to learn the concept of interest, we employ a regularization at the attention level. Specifically, we adapt the earth mover's distance (EMD) as a feature similarity metric in the cross-attention module. Moreover, benefiting from concept disentanglement, we improve the inference process and tune the prediction score by combining multiple concept probabilities. Comprehensive experiments on three CZSL benchmark datasets demonstrate that our method significantly outperforms previous works in both closed- and open-world settings, establishing a new state-of-the-art. Project page: https://haoosz.github.io/ade-czsl/
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- Beyond Seen Primitive Concepts and Attribute-Object Compositional LearningNirat Saini, Khoi Pham, Abhinav ShrivastavaCVPR 2024 · 被引用 4 次
- FlowComposer: Composable Flows for Compositional Zero-Shot LearningZhenqi He, Lin Li, Long ChenCVPR 2026 · 被引用 3 次
- Compositional Zero-shot Learning via Progressive Language-based ObservationsLin Li, Guikun Chen, Zhen Wang, Jun Xiao 等ACM MM 2025 · 被引用 2 次
- ZFusion: Efficient Deep Compositional Zero-Shot Learning for Blind Image Super-Resolution with Generative Diffusion PriorAlireza Esmaeilzehi, Hossein Zaredar, Yapeng Tian, Laleh Seyyed-KalantariICCV 2025 · 被引用 2 次
- Not Just Object, But State: Compositional Incremental Learning without ForgettingYanyi Zhang, Binglin Qiu, Qi Jia, Yu Liu 等NeurIPS 2024 · 被引用 2 次
它引用的顶会 Paper15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Do Vision Transformers See Like Convolutional Neural Networks?Maithra Raghu, Thomas Unterthiner, Simon Kornblith, Chiyuan Zhang 等NeurIPS 2021 · 被引用 1,553 次
- Task-Driven Modular Networks for Zero-Shot Compositional LearningSenthil Purushwalkam, Maximilian Nickel, Abhinav Gupta, Marc'Aurelio RanzatoICCV 2019 · 被引用 222 次
相关 Paper
- Distilled Reverse Attention Network for Open-world Compositional Zero-Shot LearningYun Li, Zhe Liu, Saurav Jha, Lina YaoICCV 2023 · 被引用 23 次
- A Conditional Probability Framework for Compositional Zero-Shot LearningPeng Wu, Qiuxia Lai, Hao Fang, Guo-Sen Xie 等ICCV 2025 · 被引用 2 次
- Learning Conditional Attributes for Compositional Zero-Shot LearningQingsheng Wang, Lingqiao Liu, Chenchen Jing, Hao Chen 等CVPR 2023
- Leveraging Sub-class Discimination for Compositional Zero-Shot LearningXiaoming Hu, Zilei WangAAAI 2023 · 被引用 21 次
- Retrieval-Augmented Primitive Representations for Compositional Zero-Shot LearningChenchen Jing, Yukun Li, Hao Chen, Chunhua ShenAAAI 2024 · 被引用 25 次
