LBMKGC: Large Model-Driven Balanced Multimodal Knowledge Graph Completion
Yuan Guo, Qian Ma, Hui Li, Qiao Ning, Furui Zhan, Yu Gu, Ge Yu, Shikai Guo
摘要
Multi-modal Knowledge Graph Completion (MMKGC) aims to predict missing entities, relations, or attributes in knowledge graphs by collaboratively modeling the triple structure and multimodal information (e.g., text, images, videos) associated with entities. This approach facilitates the automatic discovery of previously unobserved factual knowledge. However, existing MMKGC methods encounter several critical challenges: (i) the imbalance of inter-entity information across different modalities; (ii) the heterogeneity of intra-entity multimodal information; and (iii) for a given entity, the informational contributions of different modalities are inconsistent across contexts. In this paper, we propose a novel Large model-driven Balanced Multimodal Knowledge Graph Completion framework, termed LBMKGC. Initially, LBMKGC employs the Stable Diffusion XL (a Large generative vision model) to augment the imbalanced information across modalities. Subsequently, to bridge the semantic gap between heterogeneous modalities, LBMKGC aligns the multimodal embeddings of entities semantically by using the CLIP (Contrastive Language-Image Pre-Training) model. Furthermore, LBMKGC adaptively fuses multimodal embeddings with relational guidance by distinguishing between the perceptual and conceptual attributes of triples. Finally, extensive experiments conducted against 21 state-of-the-art baselines demonstrate that LBMKGC achieves superior performance across diverse datasets and scenarios while maintaining efficiency and generalizability. Our code and data are publicly available at: https://github.com/guoynow/LBMKGC.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann 等ICLR 2024 · 被引用 4,569 次
- Is Visual Context Really Helpful for Knowledge Graph? A Representation Learning PerspectiveMeng Wang, Sen Wang, Han Yang, Zheng Zhang 等ACM MM 2021 · 被引用 129 次
相关 Paper
- NativE: Multi-modal Knowledge Graph Completion in the WildYichi Zhang, Zhuo Chen, Lingbing Guo, Yajing Xu 等SIGIR 2024 · 被引用 39 次
- HFR-MKGC: Hierarchical Fusion Reasoning with MLLMs for Multi-modal Knowledge Graph CompletionDi Wang, Junping Du, Zhe Xue, Meiyu Liang 等AAAI 2026
- Multi-Modal Fact Knowledge Generation for Imbalanced Cross-Source Entity AlignmentQian Li, Cheng Ji, Zhaoji Liang, Yuzheng Zhang 等AAAI 2026
- LAFA: Multimodal Knowledge Graph Completion with Link Aware Fusion and AggregationBin Shang, Yinliang Zhao, Jun Liu, Di WangAAAI 2024 · 被引用 40 次
- Overcoming the Pitfalls of Vision-Language Model for Image-Text RetrievalFeifei Zhang, Sijia Qu, Fan Shi, Changsheng XuACM MM 2024 · 被引用 12 次
