Graph4MM: Weaving Multimodal Learning with Structural Information
Xuying Ning, Dongqi Fu, Tianxin Wei, Wujiang Xu, Jingrui He
Abstract
Real-world multimodal data usually exhibit complex structural relationships beyond traditional one-to-one mappings like image-caption pairs. Entities across modalities interact in intricate ways, with images and text forming diverse interconnections through contextual dependencies and co-references. Graphs provide powerful structural information for modeling intra-modal and intermodal relationships. However, previous works fail to distinguish multi-hop neighbors and treat the graph as a standalone modality, which fragments the overall understanding. This limitation presents two key challenges in multimodal learning: (1) integrating structural information from multi-hop neighbors into foundational models, and (2) fusing modality-specific information in a principled manner. To address these challenges, we revisit the role of graphs in multimodal learning within the era of foundation models and propose Graph4MM, a graph-based multimodal learning framework. To be specific, we introduce Hop-Diffused Attention, which integrates multi-hop structural information into selfattention through causal masking and hop diffusion. Furthermore, we design MM-QFormer, a multi-mapping querying transformer for crossmodal fusion. Through theoretical and empirical analysis, we show that leveraging structures to integrate both intra-and inter-modal interactions improves multimodal understanding beyond treating them as a standalone modality. Experiments on both generative and discriminative tasks show that Graph4MM outperforms larger VLMs, LLMs, and multimodal graph baselines, achieving a 6.93% average improvement.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9d2dcd5e-2a8f-4180-be6c-61b83f686d39Cited by top-tier papers8
- OpenMAG: A Comprehensive Benchmark for Multimodal-Attributed GraphChenxi Wan, Xunkai Li, Yilong Zuo, Haokun Deng et al.ICML 2026 · 9 citations
- Mixture of Sequence: Theme-Aware Mixture-of-Experts for Long-Sequence RecommendationXiao Lin, Zhicheng Tang, Weilin Cong, Mengyue Hang et al.WWW 2026 · 3 citations
- OwlEye: Zero-Shot Learner for Cross-Domain Graph Data Anomaly DetectionLecheng Zheng, Dongqi Fu, Zihao Li, Jingrui HeICLR 2026 · 2 citations
- Mario: Multimodal Graph Reasoning with Large Language ModelsYuanfu Sun, Kang Li, Pengkang Guo, Jiajin Liu et al.CVPR 2026 · 2 citations
- AvAtar: Learning to Align via Active Optimal TransportQi Yu, Ruizhong Qiu, Zhichen Zeng, My T. Thai et al.ICML 2026 · 1 citation
Builds on32
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNetLi Yuan, Yunpeng Chen, Tao Wang, Weihao Yu et al.ICCV 2021 · 2,462 citations
- Multimodal Few-Shot Learning with Frozen Language ModelsMaria Tsimpoukelli, Jacob Menick, Serkan Cabi, S. M. Ali Eslami et al.NeurIPS 2021 · 1,020 citations
Related papers
- A Graph Foundation Model with Cross-Modal Alignment and Modality-Aware Expert Fusion for Multi-Modal GraphsDongxiao He, AnKang Yang, Jitao Zhao, Di JinICML 2026
- Hybrid Transformer with Multi-level Fusion for Multimodal Knowledge Graph CompletionXiang Chen, Ningyu Zhang, Lei Li, Shumin Deng et al.SIGIR 2022 · 227 citations
- MOBI: Monolithic Graph-Language Modeling Beyond Modality InterferenceZhiyao Zhou, Yugang Ji, Ziwen Xu, Zhuonan Zheng et al.KDD 2026
- MolCA: Molecular Graph-Language Modeling with Cross-Modal Projector and Uni-Modal AdapterZhiyuan Liu, Sihang Li, Yanchen Luo, Hao Fei et al.EMNLP 2023 · 34 citations
- MLaGA: Multimodal Large Language and Graph AssistantDongzhe Fan, Jiajin Liu, Yi Fang, Djellel Difallah et al.KDD 2026 · 13 citations
