Multimodal Knowledge Graph Error Detection with Disentanglement VAE and Multi-Grained Triplet Confidence
Xuhui Sui, Ying Zhang, Yu Zhao, Baohang Zhou, Xiaojie Yuan
Abstract
Multimodal knowledge graphs inevitably contain numerous errors due to the lack of human supervision in their automated construction and updating processes. These errors can significantly degrade the performance of downstream applications that rely on them. Existing researches on knowledge graph error detection primarily focus on leveraging graph structural and textual information to identify triplet errors in unimodal knowledge graphs. However, unlike unimodal knowledge graphs, multimodal knowledge graphs also suffer from mismatches between images and their corresponding entities, referred to as modality errors. These modality errors not only hinder the performance of downstream applications but also impede our effective utilization of the abundant complementary information provided by the visual modality for detecting triplet errors. To this end, we introduce a novel task of multimodal knowledge graph error detection (MKGED) in this paper, aiming at simultaneously identifying both modality errors and triplet errors. Given the lack of datasets for evaluating this task, we first establish two comprehensive MKGED datasets. Furthermore, we propose a novel framework, KGDMC, to address the MKGED task. Within KGDMC, we devise a disentanglement modality reconstruction (DMR) module for modality error detection. This module disentangles each original modality representation into two disjoint components: modality-specific representations and modality-invariant representations, leveraging the cross-modality reconstruction process to detect mismatched visual modalities. Additionally, for the triplet error detection, we propose a multi-grained triplet confidence (MTC) module, incorporating local triplet confidence, global structure confidence, and global path confidence, to collaboratively detect mismatched triplets. Extensive experiments on our constructed two datasets demonstrate the superiority of our proposed framework. CCS Concepts • Computing methodologies → Knowledge representation and reasoning; Probabilistic reasoning; • Information systems → Data cleaning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- DRGW: Learning Disentangled Representations for Robust Graph WatermarkingJiasen Li, Yanwei Liu, Zhuoyi Shang, Xiaoyan Gu et al.WWW 2026 · 3 citations
- Dark Side of Modalities: Reinforced Multimodal Distillation for Multimodal Knowledge Graph ReasoningYu Zhao, Ying Zhang, Xuhui Sui, Baohang Zhou et al.ACM MM 2025
Builds on13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty et al.NeurIPS 2021 · 2,985 citations
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 2,258 citations
- Improving Multi-hop Question Answering over Knowledge Graphs using Knowledge Base EmbeddingsApoorv Saxena, Aditay Tripathi, Partha P. TalukdarACL 2020 · 488 citations
- Structure-Augmented Text Representation Learning for Efficient Knowledge Graph CompletionBo Wang, Tao Shen, Guodong Long, Tianyi Zhou et al.WWW 2021 · 322 citations
Related papers
- REMOTE: A Unified Multimodal Relation Extraction Framework with Multilevel Optimal Transport and Mixture-of-ExpertsXinkui Lin, Yongxiu Xu, Minghao Tang, Shilong Zhang et al.ACM MM 2025 · 2 citations
- Multiple Heads are Better than One: Mixture of Modality Knowledge Experts for Entity Representation LearningYichi Zhang, Zhuo Chen, Lingbing Guo, Yajing Xu et al.ICLR 2025 · 1 citation
- NativE: Multi-modal Knowledge Graph Completion in the WildYichi Zhang, Zhuo Chen, Lingbing Guo, Yajing Xu et al.SIGIR 2024 · 39 citations
- MMKGR: Multi-hop Multi-modal Knowledge Graph ReasoningShangfei Zheng, Weiqing Wang, Jianfeng Qu, Hongzhi Yin et al.ICDE 2023 · 40 citations
- VL-KGE: Vision-Language Models Meet Knowledge Graph EmbeddingsAthanasios Efthymiou, Stevan Rudinac, Monika Kackovic, Nachoem Wijnberg et al.WWW 2026 · 2 citations
