Recognizing Unseen Objects via Multimodal Intensive Knowledge Graph Propagation
Likang Wu, Zhi Li, Hongke Zhao, Zhefeng Wang, Qi Liu, Baoxing Huai, Nicholas Jing Yuan, Enhong Chen
摘要
Zero-Shot Learning (ZSL), which aims at automatically recognizing unseen objects, is a promising learning paradigm to understand new real-world knowledge for machines continuously. Recently, the Knowledge Graph (KG) has been proven as an effective scheme for handling the zero-shot task with large-scale and non-attribute data. Prior studies always embed relationships of seen and unseen objects into visual information from existing knowledge graphs to promote the cognitive ability of the unseen data. Actually, real-world knowledge is naturally formed by multimodal facts. Compared with ordinary structural knowledge from a graph perspective, multimodal KG can provide cognitive systems with fine-grained knowledge. For example, the text description and visual content can depict more critical details of a fact than only depending on knowledge triplets. Unfortunately, this multimodal fine-grained knowledge is largely unexploited due to the bottleneck of feature alignment between different modalities. To that end, we propose a multimodal intensive ZSL framework that matches regions of images with corresponding semantic embeddings via a designed dense attention module and self-calibration loss. It makes the semantic transfer process of our ZSL framework learns more differentiated knowledge between entities. Our model also gets rid of the performance limitation of only using rough global features. We conduct extensive experiments and evaluate our model on large-scale real-world data. The experimental results clearly demonstrate the effectiveness of the proposed model in standard zero-shot classification tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Hypergraph-based Zero-shot Multi-modal Product Attribute Value ExtractionJiazhen Hu, Jiaying Gong, Hongda Shen, Hoda EldardiryWWW 2025 · 被引用 4 次
- Reverse Region-to-Entity Annotation for Pixel-Level Visual Entity LinkingZhengfei Xu, Sijia Zhao, Yanchao Hao, Xiaolong Liu 等AAAI 2025
它引用的顶会 Paper10
- OntoZSL: Ontology-enhanced Zero-shot LearningYuxia Geng, Jiaoyan Chen, Zhuo Chen, Jeff Z. Pan 等WWW 2021 · 被引用 96 次
- Compositional Zero-Shot Learning via Fine-Grained Dense Feature CompositionDat Huynh, Ehsan ElhamifarNeurIPS 2020 · 被引用 89 次
- Attribute Propagation Network for Graph Zero-Shot LearningLu Liu, Tianyi Zhou, Guodong Long, Jing Jiang 等AAAI 2020 · 被引用 85 次
- Multi-modal Siamese Network for Entity AlignmentLiyi Chen, Zhi Li, Tong Xu, Han Wu 等KDD 2022 · 被引用 82 次
- VGSE: Visually-Grounded Semantic Embeddings for Zero-Shot LearningWenjia Xu, Yongqin Xian, Jiuniu Wang, Bernt Schiele 等CVPR 2022 · 被引用 61 次
相关 Paper
- Disentangled Ontology Embedding for Zero-shot LearningYuxia Geng, Jiaoyan Chen, Wen Zhang, Yajing Xu 等KDD 2022 · 被引用 22 次
- Graph Knows Unknowns: Reformulate Zero-Shot Learning as Sample-Level Graph RecognitionJingcai Guo, Song Guo, Qihua Zhou, Ziming Liu 等AAAI 2023 · 被引用 42 次
- DUET: Cross-Modal Semantic Grounding for Contrastive Zero-Shot LearningZhuo Chen, Yufeng Huang, Jiaoyan Chen, Yuxia Geng 等AAAI 2023 · 被引用 97 次
- GeoKGM: A Multimodal Large Language Model for Zero-Shot Knowledge Graph Completion in Geospatial DatabasesZhihan Zheng, Haitao Yuan, Minxiao Chen, Nan Jiang 等SIGMOD 2026 · 被引用 5 次
- Generative Adversarial Zero-Shot Relational Learning for Knowledge GraphsPengda Qin, Xin Wang, Wenhu Chen, Chunyun Zhang 等AAAI 2020 · 被引用 93 次
