Aligning Vision to Language: Annotation-Free Multimodal Knowledge Graph Construction for Enhanced LLMs Reasoning
Junming Liu, Siyuan Meng, Yanting Gao, Song Mao, Pinlong Cai, Guohang Yan, Yirong Chen, Zilin Bian, Ding Wang, Botian Shi
Abstract
Multimodal reasoning in Large Language Models (LLMs) struggles with incomplete knowledge and hallucination artifacts, challenges that textual Knowledge Graphs (KGs) only partially mitigate due to their modality isolation. While Multimodal Knowledge Graphs (MMKGs) promise enhanced cross-modal understanding, their practical construction is impeded by semantic narrowness of manual text annotations and inherent noise in visual-semantic entity linkages. In this paper, we propose Vision-align-to-Language integrated Knowledge Graph (VaLiK), a novel approach for constructing MMKGs that enhances LLMs reasoning through cross-modal information supplementation. Specifically, we cascade pre-trained Vision-Language Models (VLMs) to align image features with text, transforming them into descriptions that encapsulate image-specific information. Furthermore, we developed a cross-modal similarity verification mechanism to quantify semantic consistency, effectively filtering out noise introduced during feature alignment. Even without manually annotated image captions, the refined descriptions alone suffice to construct the MMKG. Compared to conventional MMKGs construction paradigms, our approach achieves substantial storage efficiency gains while maintaining direct entity-to-image linkage capability. Experimental results on multimodal reasoning tasks demonstrate that LLMs augmented with VaLiK outperform previous state-of-the-art models. Our code is published at https://github.com/Wings-Of-Disaster/VaLiK.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 22257990-1e01-45a8-a1f5-340878f2bdc9Cited by top-tier papers7
- HM-RAG: Hierarchical Multi-Agent Multimodal Retrieval Augmented GenerationPei Liu, Xin Liu, Ruoyu Yao, Junming Liu et al.ACM MM 2025 · 27 citations
- SceneAlign: Aligning Multimodal Reasoning to Scene Graphs in Complex Visual ScenesChuhan Wang, Xintong Li, Jennifer Yuntong Zhang, Junda Wu et al.ACL 2026 · 9 citations
- SepPrune: Structured Pruning for Efficient Deep Speech SeparationYuqi Li, Kai Li, Xin Yin, Zhifei Yang et al.AAAI 2026 · 4 citations
- Mario: Multimodal Graph Reasoning with Large Language ModelsYuanfu Sun, Kang Li, Pengkang Guo, Jiajin Liu et al.CVPR 2026 · 2 citations
- From Blind Spots to Gains: Diagnostic-Driven Iterative Training for Large Multimodal ModelsHongrui Jia, Chaoya Jiang, Yongrui Heng, Shikun Zhang et al.ICML 2026
Builds on32
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
Related papers
- Multimodal Reasoning with Multimodal Knowledge GraphJunlin Lee, Yequan Wang, Jing Li, Min ZhangACL 2024 · 29 citations
- SciMKG: A Multimodal Knowledge Graph for Science Education with Text, Image, Video and AudioTong Lu, Zhichun Wang, Yaoyu Zhou, Yiming Guan et al.AAAI 2026
- VL-KGE: Vision-Language Models Meet Knowledge Graph EmbeddingsAthanasios Efthymiou, Stevan Rudinac, Monika Kackovic, Nachoem Wijnberg et al.WWW 2026 · 2 citations
- mKG-RAG: Leveraging Multimodal Knowledge Graphs in Retrieval-Augmented Generation for Knowledge-intensive VQAXu Yuan, Liangbo Ning, Qingqing Ye, Wenqi Fan et al.SIGIR 2026 · 2 citations
- GraphVis: Boosting LLMs with Visual Knowledge Graph IntegrationYihe Deng, Chenchen Ye, Zijie Huang, Mingyu Derek Ma et al.NeurIPS 2024 · 23 citations
