Tokenization, Fusion, and Augmentation: Towards Fine-grained Multi-modal Entity Representation
Yichi Zhang, Zhuo Chen, Lingbing Guo, Yajing Xu, Binbin Hu, Ziqi Liu, Wen Zhang, Huajun Chen
Abstract
Multi-modal knowledge graph completion (MMKGC) aims to discover unobserved knowledge from given multi-modal knowledge graphs (MMKG), collaboratively leveraging structural information from the triples and multi-modal information of the entities to overcome the inherent incompleteness. Existing MMKGC methods usually extract multi-modal features with pre-trained models and employ fusion modules to integrate multi-modal features for the entities. This often results in coarse handling of multi-modal entity information, overlooking the nuanced, fine-grained semantic details and their complex interactions. To tackle this shortfall, we introduce a novel framework MyGO to tokenize, fuse, and augment the fine-grained multi-modal representations of entities and enhance the MMKGC performance. Motivated by the tokenization technology, MyGO tokenizes multi-modal entity information as fine-grained discrete tokens and learns entity representations with a cross-modal entity encoder. To further augment the multi-modal representations, MyGO incorporates fine-grained contrastive learning to highlight the specificity of the entity representations. Experiments on standard MMKGC benchmarks reveal that our method surpasses 19 of the latest models, underlining its superior performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b049f1d3-9bd4-4765-9c96-d861b08f7eb4Cited by top-tier papers5
- Mixed-Curvature Multi-Modal Knowledge Graph CompletionYuxiao Gao, Fuwei Zhang, Zhao Zhang, Xiaoshuang Min et al.AAAI 2025 · 5 citations
- LBMKGC: Large Model-Driven Balanced Multimodal Knowledge Graph CompletionYuan Guo, Qian Ma, Hui Li, Qiao Ning et al.NeurIPS 2025 · 3 citations
- Collaboration of Fusion and Independence: Hypercomplex-driven Robust Multi-Modal Knowledge Graph CompletionZhiqiang Liu, Yichi Zhang, Mengshu Sun, Lei Liang et al.ACL 2026 · 1 citation
- Dark Side of Modalities: Reinforced Multimodal Distillation for Multimodal Knowledge Graph ReasoningYu Zhao, Ying Zhang, Xuhui Sui, Baohang Zhou et al.ACM MM 2025
- HFR-MKGC: Hierarchical Fusion Reasoning with MLLMs for Multi-modal Knowledge Graph CompletionDi Wang, Junping Du, Zhe Xue, Meiyu Liang et al.AAAI 2026
Builds on15
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 2,496 citations
- TokenLearner: Adaptive Space-Time Tokenization for VideosMichael S. Ryoo, A. J. Piergiovanni, Anurag Arnab, Mostafa Dehghani et al.NeurIPS 2021 · 274 citations
- Is Visual Context Really Helpful for Knowledge Graph? A Representation Learning PerspectiveMeng Wang, Sen Wang, Han Yang, Zheng Zhang et al.ACM MM 2021 · 129 citations
- OTKGE: Multi-modal Knowledge Graph Embeddings via Optimal TransportZongsheng Cao, Qianqian Xu, Zhiyong Yang, Yuan He et al.NeurIPS 2022 · 117 citations
- IMF: Interactive Multimodal Fusion Model for Link PredictionXinhang Li, Xiangyu Zhao, Jiaxing Xu, Yong Zhang et al.WWW 2023 · 113 citations
Related papers
- Multiple Heads are Better than One: Mixture of Modality Knowledge Experts for Entity Representation LearningYichi Zhang, Zhuo Chen, Lingbing Guo, Yajing Xu et al.ICLR 2025 · 1 citation
- Contrast then Memorize: Semantic Neighbor Retrieval-Enhanced Inductive Multimodal Knowledge Graph CompletionYu Zhao, Ying Zhang, Baohang Zhou, Xinying Qian et al.SIGIR 2024 · 15 citations
- MoCoKGC: Momentum Contrast Entity Encoding for Knowledge Graph CompletionQingyang Li, Yanru Zhong, Yuchu QinEMNLP 2024 · 6 citations
- NativE: Multi-modal Knowledge Graph Completion in the WildYichi Zhang, Zhuo Chen, Lingbing Guo, Yajing Xu et al.SIGIR 2024 · 39 citations
- Improving Knowledge Graph Completion with Structure-Aware Supervised Contrastive LearningJiashi Lin, Lifang Wang, Xinyu Lu, Zhongtian Hu et al.EMNLP 2024 · 5 citations
