Tokenization, Fusion, and Augmentation: Towards Fine-grained Multi-modal Entity Representation
Yichi Zhang, Zhuo Chen, Lingbing Guo, Yajing Xu, Binbin Hu, Ziqi Liu, Wen Zhang, Huajun Chen
摘要
Multi-modal knowledge graph completion (MMKGC) aims to discover unobserved knowledge from given multi-modal knowledge graphs (MMKG), collaboratively leveraging structural information from the triples and multi-modal information of the entities to overcome the inherent incompleteness. Existing MMKGC methods usually extract multi-modal features with pre-trained models and employ fusion modules to integrate multi-modal features for the entities. This often results in coarse handling of multi-modal entity information, overlooking the nuanced, fine-grained semantic details and their complex interactions. To tackle this shortfall, we introduce a novel framework MyGO to tokenize, fuse, and augment the fine-grained multi-modal representations of entities and enhance the MMKGC performance. Motivated by the tokenization technology, MyGO tokenizes multi-modal entity information as fine-grained discrete tokens and learns entity representations with a cross-modal entity encoder. To further augment the multi-modal representations, MyGO incorporates fine-grained contrastive learning to highlight the specificity of the entity representations. Experiments on standard MMKGC benchmarks reveal that our method surpasses 19 of the latest models, underlining its superior performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Mixed-Curvature Multi-Modal Knowledge Graph CompletionYuxiao Gao, Fuwei Zhang, Zhao Zhang, Xiaoshuang Min 等AAAI 2025 · 被引用 5 次
- LBMKGC: Large Model-Driven Balanced Multimodal Knowledge Graph CompletionYuan Guo, Qian Ma, Hui Li, Qiao Ning 等NeurIPS 2025 · 被引用 3 次
- Collaboration of Fusion and Independence: Hypercomplex-driven Robust Multi-Modal Knowledge Graph CompletionZhiqiang Liu, Yichi Zhang, Mengshu Sun, Lei Liang 等ACL 2026 · 被引用 1 次
- Dark Side of Modalities: Reinforced Multimodal Distillation for Multimodal Knowledge Graph ReasoningYu Zhao, Ying Zhang, Xuhui Sui, Baohang Zhou 等ACM MM 2025
- HFR-MKGC: Hierarchical Fusion Reasoning with MLLMs for Multi-modal Knowledge Graph CompletionDi Wang, Junping Du, Zhe Xue, Meiyu Liang 等AAAI 2026
它引用的顶会 Paper15
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 被引用 2,496 次
- TokenLearner: Adaptive Space-Time Tokenization for VideosMichael S. Ryoo, A. J. Piergiovanni, Anurag Arnab, Mostafa Dehghani 等NeurIPS 2021 · 被引用 274 次
- Is Visual Context Really Helpful for Knowledge Graph? A Representation Learning PerspectiveMeng Wang, Sen Wang, Han Yang, Zheng Zhang 等ACM MM 2021 · 被引用 129 次
- OTKGE: Multi-modal Knowledge Graph Embeddings via Optimal TransportZongsheng Cao, Qianqian Xu, Zhiyong Yang, Yuan He 等NeurIPS 2022 · 被引用 117 次
- IMF: Interactive Multimodal Fusion Model for Link PredictionXinhang Li, Xiangyu Zhao, Jiaxing Xu, Yong Zhang 等WWW 2023 · 被引用 113 次
相关 Paper
- Multiple Heads are Better than One: Mixture of Modality Knowledge Experts for Entity Representation LearningYichi Zhang, Zhuo Chen, Lingbing Guo, Yajing Xu 等ICLR 2025 · 被引用 1 次
- Contrast then Memorize: Semantic Neighbor Retrieval-Enhanced Inductive Multimodal Knowledge Graph CompletionYu Zhao, Ying Zhang, Baohang Zhou, Xinying Qian 等SIGIR 2024 · 被引用 15 次
- MoCoKGC: Momentum Contrast Entity Encoding for Knowledge Graph CompletionQingyang Li, Yanru Zhong, Yuchu QinEMNLP 2024 · 被引用 6 次
- NativE: Multi-modal Knowledge Graph Completion in the WildYichi Zhang, Zhuo Chen, Lingbing Guo, Yajing Xu 等SIGIR 2024 · 被引用 39 次
- Improving Knowledge Graph Completion with Structure-Aware Supervised Contrastive LearningJiashi Lin, Lifang Wang, Xinyu Lu, Zhongtian Hu 等EMNLP 2024 · 被引用 5 次
