Beyond Entities: A Large-Scale Multi-Modal Knowledge Graph with Triplet Fact Grounding
Jingping Liu, Mingchuan Zhang, Weichen Li, Chao Wang, Shuang Li, Haiyun Jiang, Sihang Jiang, Yanghua Xiao, Yunwen Chen
摘要
Much effort has been devoted to building multi-modal knowledge graphs by visualizing entities on images, but ignoring the multi-modal information of the relation between entities. Hence, in this paper, we aim to construct a new large-scale multi-modal knowledge graph with triplet facts grounded on images that reflect not only entities but also their relations. To achieve this purpose, we propose a novel pipeline method, including triplet fact filtering, image retrieving, entity-based image filtering, relation-based image filtering, and image clustering. In this way, a multi-modal knowledge graph named ImgFact is constructed, which contains 247,732 triplet facts and 3,730,805 images. In experiments, the manual and automatic evaluations prove the reliable quality of our ImgFact. We further use the obtained images to enhance model performance on two tasks. In particular, the model optimized by our ImgFact achieves an impressive 8.38% and 9.87% improvement over the solutions enhanced by an existing multimodal knowledge graph and VisualChatGPT on F1 of relation classification. We release ImgFact and its instructions at https://github.com/kleinercubs/ImgFact.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- LLaVA Needs More Knowledge: Retrieval Augmented Natural Language Generation with Knowledge Graph for Explaining Thoracic PathologiesAmeer Hamza, Abdullah, Yong Hyun Ahn, Sungyoung Lee 等AAAI 2025 · 被引用 14 次
- I2CR: Intra- and Inter-modal Collaborative Reflections for Multimodal Entity LinkingZiyan Liu, Junwen Li, Kaiwen Li, Tong Ruan 等ACM MM 2025 · 被引用 2 次
- TeCES: Collaborative Geometric Knowledge Representation Framework under Evolving Fact SnapshotsJiujiang Guo, Zhengliang Guo, Kai Wang, Meiyang Wang 等ACL 2026
它引用的顶会 Paper7
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 被引用 2,258 次
- Learning from Context or Names? An Empirical Study on Neural Relation ExtractionHao Peng, Tianyu Gao, Xu Han, Yankai Lin 等EMNLP 2020 · 被引用 185 次
- VrR-VG: Refocusing Visually-Relevant RelationshipsYuanzhi Liang, Yalong Bai, Wei Zhang, Xueming Qian 等ICCV 2019 · 被引用 93 次
- SpatialSense: An Adversarially Crowdsourced Benchmark for Spatial Relation RecognitionKaiyu Yang, Olga Russakovsky, Jia DengICCV 2019 · 被引用 81 次
相关 Paper
- MORE: A Multimodal Object-Entity Relation Extraction Dataset with a Benchmark EvaluationLiang He, Hongke Wang, Yongchang Cao, Zhen Wu 等ACM MM 2023 · 被引用 17 次
- TIVA-KG: A Multimodal Knowledge Graph with Text, Image, Video and AudioXin Wang, Benyuan Meng, Hong Chen, Yuan Meng 等ACM MM 2023 · 被引用 71 次
- Multimodal Relation Extraction with Efficient Graph AlignmentChangmeng Zheng, Junhao Feng, Ze Fu, Yi Cai 等ACM MM 2021 · 被引用 134 次
- Multimodal Reasoning with Multimodal Knowledge GraphJunlin Lee, Yequan Wang, Jing Li, Min ZhangACL 2024 · 被引用 29 次
- LBMKGC: Large Model-Driven Balanced Multimodal Knowledge Graph CompletionYuan Guo, Qian Ma, Hui Li, Qiao Ning 等NeurIPS 2025 · 被引用 3 次
