A causal view of compositional zero-shot recognition
Yuval Atzmon, Felix Kreuk, Uri Shalit, Gal Chechik
Abstract
People easily recognize new visual categories that are new combinations of known components. This compositional generalization capacity is critical for learning in real-world domains like vision and language because the long tail of new combinations dominates the distribution. Unfortunately, learning systems struggle with compositional generalization because they often build on features that are correlated with class labels even if they are not "essential" for the class. This leads to consistent misclassification of samples from a new distribution, like new combinations of known components. Here we describe an approach for compositional generalization that builds on causal ideas. First, we describe compositional zero-shot learning from a causal perspective, and propose to view zero-shot inference as finding "which intervention caused the image?". Second, we present a causal-inspired embedding model that learns disentangled representations of elementary components of visual objects from correlated (confounded) training data. We evaluate this approach on two datasets for predicting new combinations of attribute-object pairs: A well-controlled synthesized images dataset and a real-world dataset which consists of fine-grained types of shoes. We show improvements compared to strong baselines. Code and data are provided in https://github.com/nv-research-israel/causal_comp
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d7e9183e-4b0e-4c34-bbe5-bf05b1f67f89Cited by top-tier papers40
- Linguistic Binding in Diffusion Models: Enhancing Attribute Correspondence through Attention Map AlignmentRoyi Rassin, Eran Hirsch, Daniel Glickman, Shauli Ravfogel et al.NeurIPS 2023 · 212 citations
- TransZero: Attribute-Guided Transformer for Zero-Shot LearningShiming Chen, Ziming Hong, Yang Liu, Guo-Sen Xie et al.AAAI 2022 · 185 citations
- Learning Causal Semantic Representation for Out-of-Distribution PredictionChang Liu, Xinwei Sun, Jindong Wang, Haoyue Tang et al.NeurIPS 2021 · 136 citations
- BatchFormer: Learning to Explore Sample Relationships for Robust Representation LearningZhi Hou, Baosheng Yu, Dacheng TaoCVPR 2022 · 92 citations
- Siamese Contrastive Embedding Network for Compositional Zero-Shot LearningXiangyu Li, Xu Yang, Kun Wei, Cheng Deng et al.CVPR 2022 · 87 citations
Builds on4
- Task-Driven Modular Networks for Zero-Shot Compositional LearningSenthil Purushwalkam, Maximilian Nickel, Abhinav Gupta, Marc'Aurelio RanzatoICCV 2019 · 222 citations
- Adversarial Fine-Grained Composition Learning for Unseen Attribute-Object RecognitionKun Wei, Muli Yang, Hao Wang, Cheng Deng et al.ICCV 2019 · 95 citations
- Robust Learning with the Hilbert-Schmidt Independence CriterionDaniel Greenfeld, Uri ShalitICML 2020 · 73 citations
- Symmetry and Group in Attribute-Object CompositionsYong-Lu Li, Yue Xu, Xiaohan Mao, Cewu LuCVPR 2020
Related papers
- Learning Conditional Attributes for Compositional Zero-Shot LearningQingsheng Wang, Lingqiao Liu, Chenchen Jing, Hao Chen et al.CVPR 2023
- Learning Attention as Disentangler for Compositional Zero-Shot LearningShaozhe Hao, Kai Han, Kwan-Yee K. WongCVPR 2023
- I2CD: An Invertible Causal Framework for Compositional Zero-Shot Learning via Disentangle-Compose-DisentangleZhaoquan Yuan, Zining Wang, Yuankang Pan, Ao Luo et al.AAAI 2026
- Independent Prototype Propagation for Zero-Shot CompositionalityFrank Ruis, Gertjan J. Burghouts, Doina BucurNeurIPS 2021 · 77 citations
- Disentangling Visual Embeddings for Attributes and ObjectsNirat Saini, Khoi Pham, Abhinav ShrivastavaCVPR 2022 · 74 citations
