Recognizing Unseen Objects via Multimodal Intensive Knowledge Graph Propagation
Likang Wu, Zhi Li, Hongke Zhao, Zhefeng Wang, Qi Liu, Baoxing Huai, Nicholas Jing Yuan, Enhong Chen
Abstract
Zero-Shot Learning (ZSL), which aims at automatically recognizing unseen objects, is a promising learning paradigm to understand new real-world knowledge for machines continuously. Recently, the Knowledge Graph (KG) has been proven as an effective scheme for handling the zero-shot task with large-scale and non-attribute data. Prior studies always embed relationships of seen and unseen objects into visual information from existing knowledge graphs to promote the cognitive ability of the unseen data. Actually, real-world knowledge is naturally formed by multimodal facts. Compared with ordinary structural knowledge from a graph perspective, multimodal KG can provide cognitive systems with fine-grained knowledge. For example, the text description and visual content can depict more critical details of a fact than only depending on knowledge triplets. Unfortunately, this multimodal fine-grained knowledge is largely unexploited due to the bottleneck of feature alignment between different modalities. To that end, we propose a multimodal intensive ZSL framework that matches regions of images with corresponding semantic embeddings via a designed dense attention module and self-calibration loss. It makes the semantic transfer process of our ZSL framework learns more differentiated knowledge between entities. Our model also gets rid of the performance limitation of only using rough global features. We conduct extensive experiments and evaluate our model on large-scale real-world data. The experimental results clearly demonstrate the effectiveness of the proposed model in standard zero-shot classification tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 01feb213-cecf-4e90-beb6-d42dd6c6374fCited by top-tier papers2
- Hypergraph-based Zero-shot Multi-modal Product Attribute Value ExtractionJiazhen Hu, Jiaying Gong, Hongda Shen, Hoda EldardiryWWW 2025 · 4 citations
- Reverse Region-to-Entity Annotation for Pixel-Level Visual Entity LinkingZhengfei Xu, Sijia Zhao, Yanchao Hao, Xiaolong Liu et al.AAAI 2025
Builds on10
- OntoZSL: Ontology-enhanced Zero-shot LearningYuxia Geng, Jiaoyan Chen, Zhuo Chen, Jeff Z. Pan et al.WWW 2021 · 96 citations
- Compositional Zero-Shot Learning via Fine-Grained Dense Feature CompositionDat Huynh, Ehsan ElhamifarNeurIPS 2020 · 89 citations
- Attribute Propagation Network for Graph Zero-Shot LearningLu Liu, Tianyi Zhou, Guodong Long, Jing Jiang et al.AAAI 2020 · 85 citations
- Multi-modal Siamese Network for Entity AlignmentLiyi Chen, Zhi Li, Tong Xu, Han Wu et al.KDD 2022 · 82 citations
- VGSE: Visually-Grounded Semantic Embeddings for Zero-Shot LearningWenjia Xu, Yongqin Xian, Jiuniu Wang, Bernt Schiele et al.CVPR 2022 · 61 citations
Related papers
- Disentangled Ontology Embedding for Zero-shot LearningYuxia Geng, Jiaoyan Chen, Wen Zhang, Yajing Xu et al.KDD 2022 · 22 citations
- Graph Knows Unknowns: Reformulate Zero-Shot Learning as Sample-Level Graph RecognitionJingcai Guo, Song Guo, Qihua Zhou, Ziming Liu et al.AAAI 2023 · 42 citations
- DUET: Cross-Modal Semantic Grounding for Contrastive Zero-Shot LearningZhuo Chen, Yufeng Huang, Jiaoyan Chen, Yuxia Geng et al.AAAI 2023 · 97 citations
- GeoKGM: A Multimodal Large Language Model for Zero-Shot Knowledge Graph Completion in Geospatial DatabasesZhihan Zheng, Haitao Yuan, Minxiao Chen, Nan Jiang et al.SIGMOD 2026 · 5 citations
- Generative Adversarial Zero-Shot Relational Learning for Knowledge GraphsPengda Qin, Xin Wang, Wenhu Chen, Chunyun Zhang et al.AAAI 2020 · 93 citations
