Hierarchical Visual Primitive Experts for Compositional Zero-Shot Learning
Hanjae Kim, Jiyoung Lee, Seongheon Park, Kwanghoon Sohn
摘要
Compositional zero-shot learning (CZSL) aims to recognize unseen compositions with prior knowledge of known primitives (attribute and object). Previous works for CZSL often suffer from grasping the contextuality between attribute and object, as well as the discriminability of visual features, and the long-tailed distribution of real-world compositional data. We propose a simple and scalable framework called Composition Transformer (CoT) to address these issues. CoT employs object and attribute experts in distinctive manners to generate representative embeddings, using the visual network hierarchically. The object expert extracts representative object embeddings from the final layer in a bottom-up manner, while the attribute expert makes attribute embeddings in a top-down manner with a proposed object-guided attention module that models contextuality explicitly. To remedy biased prediction caused by imbalanced data distribution, we develop a simple minority attribute augmentation (MAA) that synthesizes virtual samples by mixing two images and oversampling minority attribute classes. Our method achieves SoTA performance on several benchmarks, including MIT-States, C-GQA, and VAW-CZSL. We also demonstrate the effectiveness of CoT in improving visual discrimination and addressing the model bias from the imbalanced data distribution. The code is available at https://github.com/HanjaeKim98/ CoT .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- FlowComposer: Composable Flows for Compositional Zero-Shot LearningZhenqi He, Lin Li, Long ChenCVPR 2026 · 被引用 3 次
- Compositional Zero-shot Learning via Progressive Language-based ObservationsLin Li, Guikun Chen, Zhen Wang, Jun Xiao 等ACM MM 2025 · 被引用 2 次
- Not Just Object, But State: Compositional Incremental Learning without ForgettingYanyi Zhang, Binglin Qiu, Qi Jia, Yu Liu 等NeurIPS 2024 · 被引用 2 次
- A Conditional Probability Framework for Compositional Zero-Shot LearningPeng Wu, Qiuxia Lai, Hao Fang, Guo-Sen Xie 等ICCV 2025 · 被引用 2 次
- Zero-Shot Compositional Video Learning with Coding Rate ReductionHeeseok Jung, Jun-Hyeon Bak, Yujin Jeong, Gyugeun Lee 等ICCV 2025 · 被引用 1 次
它引用的顶会 Paper30
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh 等ICCV 2019 · 被引用 5,843 次
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 被引用 2,647 次
- Long-Tailed Classification by Keeping the Good and Removing the Bad Momentum Causal EffectKaihua Tang, Jianqiang Huang, Hanwang ZhangNeurIPS 2020 · 被引用 533 次
相关 Paper
- Revealing the Proximate Long-Tail Distribution in Compositional Zero-Shot LearningChenyi Jiang, Haofeng ZhangAAAI 2024 · 被引用 18 次
- Learning Conditional Attributes for Compositional Zero-Shot LearningQingsheng Wang, Lingqiao Liu, Chenchen Jing, Hao Chen 等CVPR 2023
- Learning Graph Embeddings for Compositional Zero-Shot LearningMuhammad Ferjad Naeem, Yongqin Xian, Federico Tombari, Zeynep AkataCVPR 2021
- Decoupling Primitive with Experts: Dynamic Feature Alignment for Compositional Zero-Shot LearningXiao Zhang, Haodong Jing, Yongqiang MA, Nanning ZhengICLR 2026
- Retrieval-Augmented Primitive Representations for Compositional Zero-Shot LearningChenchen Jing, Yukun Li, Hao Chen, Chunhua ShenAAAI 2024 · 被引用 25 次
