rMMEA: Robust Multi-Modal Entity Alignment with Missing and Noise Visual Modality
Lingbing Guo, Zhuo Chen, Yichi Zhang, Wenbin Guo, Haonan Yang, Zhao Li, Zirui Chen, Xin Wang
Abstract
Recently, multi-modal embedding methods have flourished in entity alignment. As state-of-the-art approaches evolve rapidly, visual modality (i.e., images) missing emerges as a critical challenge. While visual modality typically offers the most informative signals in multi-modal entity alignment (MMEA), it is frequently unavailable for many entities. The existing methods commonly use dummy vectors to represent visual-missing embeddings, which negatively impacts both model training and inference. In this paper, we propose robust multi-modal entity alignment (rMMEA), which leverages ranking-based knowledge distillation and mutual information (MI) estimation to address missing modalities while enhancing noise robustness. Unlike conventional teacher-student distillation that requires the student to replicate teacher outputs, our rMMEA learns soft rankings from pure and complete modality sides while capturing implicit key semantics of teacher embeddings through mutual information maximization, allowing rMMEA to avoid strict point-to-point alignment. The experimental results across multiple benchmarks and settings demonstrate that rMMEA significantly outperforms the state-of-the-art anti-modality-missing methods in terms of effectiveness and efficiency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 08b6d017-a3c8-4897-ba5b-31e4cb54ed6aBuilds on16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Contrastive Representation DistillationYonglong Tian, Dilip Krishnan, Phillip IsolaICLR 2020 · 1,305 citations
- Knowledge Graph Alignment Network with Gated Multi-Hop Neighborhood AggregationZequn Sun, Chengming Wang, Wei Hu, Muhao Chen et al.AAAI 2020 · 379 citations
- Visual Pivoting for (Unsupervised) Entity AlignmentFangyu Liu, Muhao Chen, Dan Roth, Nigel CollierAAAI 2021 · 159 citations
Related papers
- On Modality Weighting and Specificity for Multi-Modal Entity AlignmentYu Xing, Qizhuo Xie, Yunhui Liu, Qing Gu et al.AAAI 2026
- Explicit-Implicit Entity Alignment Method in Multi-modal Knowledge GraphsLuyao Wang, Chunlai Zhou, Biao QinKDD 2025
- Boosting Discriminability for Robust Multimodal Entity Linking with Visual Modality MissingMingrui Lao, Zheng Li, Yanming Guo, Xueyi Zhang et al.SIGIR 2025 · 4 citations
- Cross-Modal Graph Attention Network for Entity AlignmentBaogui Xu, Chengjin Xu, Bing SuACM MM 2023 · 21 citations
- Mitigating Modality Bias in Multi-modal Entity Alignment from a Causal PerspectiveTaoyu Su, Jiawei Sheng, Duohe Ma, Xiaodong Li et al.SIGIR 2025 · 4 citations
