Mosaic of Modalities: A Comprehensive Benchmark for Multimodal Graph Learning
Jing Zhu, Yuhang Zhou, Shengyi Qian, Zhongmou He, Tong Zhao, Neil Shah, Danai Koutra
2025年份
9顶会引用
摘要
Figure 1. Visualization of our Multimodal Graph Benchmark (MM-GRAPH). All nodes of our benchmark have both visual and text features. (a) Amazon-Sports: The image and text come from the original image and title of the sports equipment. (b) Goodreads-LP: The images correspond to book covers. We do not show the text features of Goodreads-LP since the book description is very long. (c) Ele-fashion: The images and texts correspond to the original image and title of the fashion product, respectively.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- MLaGA: Multimodal Large Language and Graph AssistantDongzhe Fan, Jiajin Liu, Yi Fang, Djellel Difallah 等KDD 2026 · 被引用 13 次
- OpenMAG: A Comprehensive Benchmark for Multimodal-Attributed GraphChenxi Wan, Xunkai Li, Yilong Zuo, Haokun Deng 等ICML 2026 · 被引用 9 次
- Multimodal Graph Representation Learning with Dynamic Information PathwaysXiaobin Hong, Mingkai Lin, Xiaoli Wang, Chaoqun Wang 等AAAI 2026 · 被引用 1 次
- Decoupling and Damping: Structurally-Regularized Gradient Matching for Multimodal Graph CondensationLian Shen, Zhendan Chen, Meijia Song, Yinghui Jiang 等KDD 2026 · 被引用 1 次
- GraphVLM: Benchmarking Vision Language Models for Multimodal Graph LearningJiajin Liu, Dongzhe Fan, Chuanhao Ji, Daochen Zha 等CVPR 2026 · 被引用 1 次
它引用的顶会 Paper16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Open Graph Benchmark: Datasets for Machine Learning on GraphsWeihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong 等NeurIPS 2020 · 被引用 3,935 次
- Perceiver: General Perception with Iterative AttentionAndrew Jaegle, Felix Gimeno, Andy Brock, Oriol Vinyals 等ICML 2021 · 被引用 1,399 次
- FLAVA: A Foundational Language And Vision Alignment ModelAmanpreet Singh, Ronghang Hu, Vedanuj Goswami, Guillaume Couairon 等CVPR 2022 · 被引用 483 次
相关 Paper
- Fashion Retrieval via Graph Reasoning Networks on a Similarity PyramidZhanghui Kuang, Yiming Gao, Guanbin Li, Ping Luo 等ICCV 2019 · 被引用 105 次
- Hypergraph-based Zero-shot Multi-modal Product Attribute Value ExtractionJiazhen Hu, Jiaying Gong, Hongda Shen, Hoda EldardiryWWW 2025 · 被引用 4 次
- Multi-modal Multi-relational Feature Aggregation Network for Medical Knowledge Representation LearningYingying Zhang, Quan Fang, Shengsheng Qian, Changsheng XuACM MM 2020 · 被引用 20 次
- FaD-VLP: Fashion Vision-and-Language Pre-training towards Unified Retrieval and CaptioningSuvir Mirchandani, Licheng Yu, Mengjiao Wang, Animesh Sinha 等EMNLP 2022 · 被引用 9 次
- Personalized Fashion Compatibility Modeling via Metapath-guided Heterogeneous Graph LearningWeili Guan, Fangkai Jiao, Xuemeng Song, Haokun Wen 等SIGIR 2022 · 被引用 51 次
