HFR-MKGC: Hierarchical Fusion Reasoning with MLLMs for Multi-modal Knowledge Graph Completion
Di Wang, Junping Du, Zhe Xue, Meiyu Liang, Guanhua Ye, Yingxia Shao, Haisheng Li
摘要
Multi-modal knowledge graph completion (MMKGC) aims to infer missing entities of triples by leveraging heterogeneous information in knowledge graph (KG). However, existing approaches often struggle with inconsistent modality alignment, limited reasoning depth, and insufficient negative sample quality. In this work, we propose HFR-MKGC, a novel framework that integrates hierarchical modal fusion and Multimodal Large Language Model (MLLM) reasoning for robust and expressive MMKGC. Specifically, we introduce a relation-guided hierarchical modal fusion module, which conducts fine-grained intra-visual fusion and relation-guided cross-modal integration to yield rich entity representations. HFR-MKGC employs a fine-tuned MLLM to perform instruction-based triple reasoning, producing candidate entities for completion. Then, it constructs hard negative samples through textual perturbation by MLLM and visual feature augmentation with rotation and noise. HFR-MKGC optimizes the model via adversarial training. Extensive experiments on three MMKGC benchmarks demonstrate that our method outperforms state-of-the-art methods, validating its effectiveness in MMKGC.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Hybrid Transformer with Multi-level Fusion for Multimodal Knowledge Graph CompletionXiang Chen, Ningyu Zhang, Lei Li, Shumin Deng 等SIGIR 2022 · 被引用 227 次
- Is Visual Context Really Helpful for Knowledge Graph? A Representation Learning PerspectiveMeng Wang, Sen Wang, Han Yang, Zheng Zhang 等ACM MM 2021 · 被引用 129 次
- Relation-enhanced Negative Sampling for Multimodal Knowledge Graph CompletionDerong Xu, Tong Xu, Shiwei Wu, Jingbo Zhou 等ACM MM 2022 · 被引用 89 次
相关 Paper
- NativE: Multi-modal Knowledge Graph Completion in the WildYichi Zhang, Zhuo Chen, Lingbing Guo, Yajing Xu 等SIGIR 2024 · 被引用 39 次
- LBMKGC: Large Model-Driven Balanced Multimodal Knowledge Graph CompletionYuan Guo, Qian Ma, Hui Li, Qiao Ning 等NeurIPS 2025 · 被引用 3 次
- Multiple Heads are Better than One: Mixture of Modality Knowledge Experts for Entity Representation LearningYichi Zhang, Zhuo Chen, Lingbing Guo, Yajing Xu 等ICLR 2025 · 被引用 1 次
- Multi-Modal Fact Knowledge Generation for Imbalanced Cross-Source Entity AlignmentQian Li, Cheng Ji, Zhaoji Liang, Yuzheng Zhang 等AAAI 2026
- LAFA: Multimodal Knowledge Graph Completion with Link Aware Fusion and AggregationBin Shang, Yinliang Zhao, Jun Liu, Di WangAAAI 2024 · 被引用 40 次
