On Modality Weighting and Specificity for Multi-Modal Entity Alignment
Yu Xing, Qizhuo Xie, Yunhui Liu, Qing Gu, Tao Zheng, Bin Chong, Tieke He
Abstract
Multi-modal entity alignment aims to identify equivalent entities across different multi-modal knowledge graphs (MMKGs). While prior work has achieved notable progress through improved multi-modal encoding and cross-modal fusion techniques, two critical challenges remain unresolved. First, due to the heterogeneous and often inconsistent sources from which MMKGs are constructed, the quality and informativeness of modalities vary significantly across entities, leading to the modality weighting problem. Second, existing cross-modal fusion mechanisms predominantly emphasize modality-shared information, often at the expense of modality-specific signals that are also essential for precise alignment. To address these issues, we propose HUMEA, a novel framework that integrates hierarchical Mixture-of-Experts (MoE) with unimodal distillation. HUMEA consists of: (1) A hierarchical MoE module comprising intra-modal and inter-modal experts, which adaptively modulates modality contributions by capturing entity representations at fine-to-coarse semantic granularities. In addition, we introduce a contrastive mutual information loss to enhance expert diversity and reduce redundancy. (2) A unimodal distillation strategy that preserves modality-specific information in the fused representations through single-modality alignment and distillation, achieving a balanced integration of shared and unique modality features. Extensive experiments on two benchmark datasets, FB15K-DB15K and FB15K-YAGO15K, demonstrate state-of-the-art performance, validating the effectiveness of our approach.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b7de90ed-8554-4d8f-a932-0420d140fdb3Builds on19
- CLUB: A Contrastive Log-ratio Upper Bound of Mutual InformationPengyu Cheng, Weituo Hao, Shuyang Dai, Jiachang Liu et al.ICML 2020 · 512 citations
- Towards Universal Sequence Representation Learning for Recommender SystemsYupeng Hou, Shanlei Mu, Wayne Xin Zhao, Yaliang Li et al.KDD 2022 · 245 citations
- LGMRec: Local and Global Graph Learning for Multimodal RecommendationZhiqiang Guo, Jianjun Li, Guohui Li, Chaoyang Wang et al.AAAI 2024 · 164 citations
- Visual Pivoting for (Unsupervised) Entity AlignmentFangyu Liu, Muhao Chen, Dan Roth, Nigel CollierAAAI 2021 · 159 citations
- FuseMoE: Mixture-of-Experts Transformers for Fleximodal FusionXing Han, Huy Nguyen, Carl Harris, Nhat Ho et al.NeurIPS 2024 · 129 citations
Related papers
- Enhancing Multi-Modal Entity Alignment via Multi-Grained Decision FusionYu Xing, Qizhuo Xie, You Lv, Ziyang Zhou et al.WWW 2026
- IBMEA: Exploring Variational Information Bottleneck for Multi-modal Entity AlignmentTaoyu Su, Jiawei Sheng, Shicheng Wang, Xinghua Zhang et al.ACM MM 2024 · 7 citations
- Multi-modal Siamese Network for Entity AlignmentLiyi Chen, Zhi Li, Tong Xu, Han Wu et al.KDD 2022 · 82 citations
- Attribute-Consistent Knowledge Graph Representation Learning for Multi-Modal Entity AlignmentQian Li, Shu Guo, Yangyifei Luo, Cheng Ji et al.WWW 2023 · 56 citations
- MEAformer: Multi-modal Entity Alignment Transformer for Meta Modality HybridZhuo Chen, Jiaoyan Chen, Wen Zhang, Lingbing Guo et al.ACM MM 2023 · 66 citations
