Mosaic of Modalities: A Comprehensive Benchmark for Multimodal Graph Learning
Jing Zhu, Yuhang Zhou, Shengyi Qian, Zhongmou He, Tong Zhao, Neil Shah, Danai Koutra
2025Year
9Top-tier citations
Abstract
Figure 1. Visualization of our Multimodal Graph Benchmark (MM-GRAPH). All nodes of our benchmark have both visual and text features. (a) Amazon-Sports: The image and text come from the original image and title of the sports equipment. (b) Goodreads-LP: The images correspond to book covers. We do not show the text features of Goodreads-LP since the book description is very long. (c) Ele-fashion: The images and texts correspond to the original image and title of the fashion product, respectively.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- MLaGA: Multimodal Large Language and Graph AssistantDongzhe Fan, Jiajin Liu, Yi Fang, Djellel Difallah et al.KDD 2026 · 13 citations
- OpenMAG: A Comprehensive Benchmark for Multimodal-Attributed GraphChenxi Wan, Xunkai Li, Yilong Zuo, Haokun Deng et al.ICML 2026 · 9 citations
- Multimodal Graph Representation Learning with Dynamic Information PathwaysXiaobin Hong, Mingkai Lin, Xiaoli Wang, Chaoqun Wang et al.AAAI 2026 · 1 citation
- Decoupling and Damping: Structurally-Regularized Gradient Matching for Multimodal Graph CondensationLian Shen, Zhendan Chen, Meijia Song, Yinghui Jiang et al.KDD 2026 · 1 citation
- GraphVLM: Benchmarking Vision Language Models for Multimodal Graph LearningJiajin Liu, Dongzhe Fan, Chuanhao Ji, Daochen Zha et al.CVPR 2026 · 1 citation
Builds on16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Open Graph Benchmark: Datasets for Machine Learning on GraphsWeihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong et al.NeurIPS 2020 · 3,935 citations
- Perceiver: General Perception with Iterative AttentionAndrew Jaegle, Felix Gimeno, Andy Brock, Oriol Vinyals et al.ICML 2021 · 1,399 citations
- FLAVA: A Foundational Language And Vision Alignment ModelAmanpreet Singh, Ronghang Hu, Vedanuj Goswami, Guillaume Couairon et al.CVPR 2022 · 483 citations
Related papers
- Fashion Retrieval via Graph Reasoning Networks on a Similarity PyramidZhanghui Kuang, Yiming Gao, Guanbin Li, Ping Luo et al.ICCV 2019 · 105 citations
- Hypergraph-based Zero-shot Multi-modal Product Attribute Value ExtractionJiazhen Hu, Jiaying Gong, Hongda Shen, Hoda EldardiryWWW 2025 · 4 citations
- Multi-modal Multi-relational Feature Aggregation Network for Medical Knowledge Representation LearningYingying Zhang, Quan Fang, Shengsheng Qian, Changsheng XuACM MM 2020 · 20 citations
- FaD-VLP: Fashion Vision-and-Language Pre-training towards Unified Retrieval and CaptioningSuvir Mirchandani, Licheng Yu, Mengjiao Wang, Animesh Sinha et al.EMNLP 2022 · 9 citations
- Personalized Fashion Compatibility Modeling via Metapath-guided Heterogeneous Graph LearningWeili Guan, Fangkai Jiao, Xuemeng Song, Haokun Wen et al.SIGIR 2022 · 51 citations
