TIVA-KG: A Multimodal Knowledge Graph with Text, Image, Video and Audio
Xin Wang, Benyuan Meng, Hong Chen, Yuan Meng, Ke Lv, Wenwu Zhu
Abstract
Knowledge graphs serve as a powerful tool to boost model performances for various applications covering computer vision, natural language processing, multimedia data mining, etc. The process of knowledge acquisition for human is multimodal in essence, covering text, image, video and audio modalities. However, existing multimodal knowledge graphs fail to cover all these four elements simultaneously, severely limiting their expressive powers in performance improvement for downstream tasks. In this paper, we propose TIVA-KG, a multimodal Knowledge Graph covering Text, Image, Video and Audio, which can benefit various downstream tasks. Our proposed TIVA-KG has two significant advantages over existing knowledge graphs in i) coverage of up to four modalities including text, image, video, audio, and ii) capability of triplet grounding which grounds multimodal relations to triples instead of entities. We further design a Quadruple Embedding Baseline (QEB) model to validate the necessity and efficacy of considering four modalities in KG. We conduct extensive experiments to test the proposed TIVA-KG with various knowledge graph representation approaches over link prediction task, demonstrating the benefits and necessity of introducing multiple modalities and triplet grounding. TIVA-KG is expected to promote further research on mining multimodal knowledge graph as well as the relevant downstream tasks in the community. TIVA-KG is now available at our website: http://mn.cs.tsinghua.edu.cn/tivakg.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers20
- NativE: Multi-modal Knowledge Graph Completion in the WildYichi Zhang, Zhuo Chen, Lingbing Guo, Yajing Xu et al.SIGIR 2024 · 39 citations
- Aligning Vision to Language: Annotation-Free Multimodal Knowledge Graph Construction for Enhanced LLMs ReasoningJunming Liu, Siyuan Meng, Yanting Gao, Song Mao et al.ICCV 2025 · 34 citations
- Tokenization, Fusion, and Augmentation: Towards Fine-grained Multi-modal Entity RepresentationYichi Zhang, Zhuo Chen, Lingbing Guo, Yajing Xu et al.AAAI 2025 · 25 citations
- APKGC: Noise-enhanced Multi-Modal Knowledge Graph Completion with Attention PenaltyYue Jian, Xiangyu Luo, Zhifei Li, Miao Zhang et al.AAAI 2025 · 21 citations
- Test-Time Training on Graphs with Large Language Models (LLMs)Jiaxin Zhang, Yiqi Wang, Xihong Yang, Siwei Wang et al.ACM MM 2024 · 6 citations
Builds on5
- RNNLogic: Learning Logic Rules for Reasoning on Knowledge GraphsMeng Qu, Jun-Kun Chen, Louis-Pascal A. C. Xhonneux, Yoshua Bengio et al.ICLR 2021 · 230 citations
- Boosting Visual Question Answering with Context-aware Knowledge AggregationGuohao Li, Xin Wang, Wenwu ZhuACM MM 2020 · 82 citations
- AVQA: A Dataset for Audio-Visual Question Answering on VideosPinci Yang, Xin Wang, Xuguang Duan, Hong Chen et al.ACM MM 2022 · 60 citations
- From Strings to Things: Knowledge-Enabled VQA Model That Can Read and ReasonAjeet Kumar Singh, Anand Mishra, Shashank Shekhar, Anirban ChakrabortyICCV 2019 · 54 citations
- Hierarchical Conditional Relation Networks for Video Question AnsweringThao Minh Le, Vuong Le, Svetha Venkatesh, Truyen TranCVPR 2020
Related papers
- WFF: Wavelet-based Information Fusion for Multimodal Knowledge Graph Link PredictionXiaodi Xu, Lijie Li, Ye Wang, Tao Ren et al.ACM MM 2025 · 1 citation
- VL-KGE: Vision-Language Models Meet Knowledge Graph EmbeddingsAthanasios Efthymiou, Stevan Rudinac, Monika Kackovic, Nachoem Wijnberg et al.WWW 2026 · 2 citations
- Hybrid Transformer with Multi-level Fusion for Multimodal Knowledge Graph CompletionXiang Chen, Ningyu Zhang, Lei Li, Shumin Deng et al.SIGIR 2022 · 227 citations
- Multimodal Biological Knowledge Graph Completion via Triple Co-Attention MechanismDerong Xu, Jingbo Zhou, Tong Xu, Yuan Xia et al.ICDE 2023 · 24 citations
- InGram: Inductive Knowledge Graph Embedding via Relation GraphsJaejun Lee, Chanyoung Chung, Joyce Jiyoung WhangICML 2023 · 83 citations
