Disentangling Visual Embeddings for Attributes and Objects
Nirat Saini, Khoi Pham, Abhinav Shrivastava
Abstract
We study the problem of compositional zero-shot learning for object-attribute recognition. Prior works use visual features extracted with a backbone network, pre-trained for object classification and thus do not capture the subtly distinct features associated with attributes. To overcome this challenge, these studies employ supervision from the linguistic space, and use pre-trained word embeddings to better separate and compose attribute-object pairs for recognition. Analogous to linguistic embedding space, which already has unique and agnostic embeddings for object and attribute, we shift the focus back to the visual space and propose a novel architecture that can disentangle attribute and object features in the visual space. We use visual decomposed features to hallucinate embeddings that are representative for the seen and novel compositions to better regularize the learning of our model. Extensive experiments show that our method outperforms existing work with significant margin on three datasets: MIT-States, UT-Zappos, and a new benchmark created based on VAW. The code, models, and dataset splits are publicly available at https: //github.com/nirat1606/OADis .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e25d0e8f-e015-44da-b924-07b9adb142efCited by top-tier papers26
- Composing Object Relations and Attributes for Image-Text MatchingKhoi Pham, Chuong Huynh, Ser-Nam Lim, Abhinav ShrivastavaCVPR 2024 · 31 citations
- Hierarchical Visual Primitive Experts for Compositional Zero-Shot LearningHanjae Kim, Jiyoung Lee, Seongheon Park, Kwanghoon SohnICCV 2023 · 27 citations
- Retrieval-Augmented Primitive Representations for Compositional Zero-Shot LearningChenchen Jing, Yukun Li, Hao Chen, Chunhua ShenAAAI 2024 · 25 citations
- Distilled Reverse Attention Network for Open-world Compositional Zero-Shot LearningYun Li, Zhe Liu, Saurav Jha, Lina YaoICCV 2023 · 23 citations
- Leveraging Sub-class Discimination for Compositional Zero-Shot LearningXiaoming Hu, Zilei WangAAAI 2023 · 21 citations
Builds on14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- Visual Semantic Reasoning for Image-Text MatchingKunpeng Li, Yulun Zhang, Kai Li, Yuanyuan Li et al.ICCV 2019 · 598 citations
- Task-Driven Modular Networks for Zero-Shot Compositional LearningSenthil Purushwalkam, Maximilian Nickel, Abhinav Gupta, Marc'Aurelio RanzatoICCV 2019 · 222 citations
- A causal view of compositional zero-shot recognitionYuval Atzmon, Felix Kreuk, Uri Shalit, Gal ChechikNeurIPS 2020 · 163 citations
Related papers
- Learning Conditional Attributes for Compositional Zero-Shot LearningQingsheng Wang, Lingqiao Liu, Chenchen Jing, Hao Chen et al.CVPR 2023
- Beyond Seen Primitive Concepts and Attribute-Object Compositional LearningNirat Saini, Khoi Pham, Abhinav ShrivastavaCVPR 2024 · 4 citations
- VGSE: Visually-Grounded Semantic Embeddings for Zero-Shot LearningWenjia Xu, Yongqin Xian, Jiuniu Wang, Bernt Schiele et al.CVPR 2022 · 61 citations
- Learning Attention as Disentangler for Compositional Zero-Shot LearningShaozhe Hao, Kai Han, Kwan-Yee K. WongCVPR 2023
- Learning Visual Proxy for Compositional Zero-Shot LearningShiyu Zhang, Cheng Yan, Yang Liu, Chenchen Jing et al.ICCV 2025 · 1 citation
