TIVA-KG: A Multimodal Knowledge Graph with Text, Image, Video and Audio
Xin Wang, Benyuan Meng, Hong Chen, Yuan Meng, Ke Lv, Wenwu Zhu
摘要
Knowledge graphs serve as a powerful tool to boost model performances for various applications covering computer vision, natural language processing, multimedia data mining, etc. The process of knowledge acquisition for human is multimodal in essence, covering text, image, video and audio modalities. However, existing multimodal knowledge graphs fail to cover all these four elements simultaneously, severely limiting their expressive powers in performance improvement for downstream tasks. In this paper, we propose TIVA-KG, a multimodal Knowledge Graph covering Text, Image, Video and Audio, which can benefit various downstream tasks. Our proposed TIVA-KG has two significant advantages over existing knowledge graphs in i) coverage of up to four modalities including text, image, video, audio, and ii) capability of triplet grounding which grounds multimodal relations to triples instead of entities. We further design a Quadruple Embedding Baseline (QEB) model to validate the necessity and efficacy of considering four modalities in KG. We conduct extensive experiments to test the proposed TIVA-KG with various knowledge graph representation approaches over link prediction task, demonstrating the benefits and necessity of introducing multiple modalities and triplet grounding. TIVA-KG is expected to promote further research on mining multimodal knowledge graph as well as the relevant downstream tasks in the community. TIVA-KG is now available at our website: http://mn.cs.tsinghua.edu.cn/tivakg.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper20
- NativE: Multi-modal Knowledge Graph Completion in the WildYichi Zhang, Zhuo Chen, Lingbing Guo, Yajing Xu 等SIGIR 2024 · 被引用 39 次
- Aligning Vision to Language: Annotation-Free Multimodal Knowledge Graph Construction for Enhanced LLMs ReasoningJunming Liu, Siyuan Meng, Yanting Gao, Song Mao 等ICCV 2025 · 被引用 34 次
- Tokenization, Fusion, and Augmentation: Towards Fine-grained Multi-modal Entity RepresentationYichi Zhang, Zhuo Chen, Lingbing Guo, Yajing Xu 等AAAI 2025 · 被引用 25 次
- APKGC: Noise-enhanced Multi-Modal Knowledge Graph Completion with Attention PenaltyYue Jian, Xiangyu Luo, Zhifei Li, Miao Zhang 等AAAI 2025 · 被引用 21 次
- Test-Time Training on Graphs with Large Language Models (LLMs)Jiaxin Zhang, Yiqi Wang, Xihong Yang, Siwei Wang 等ACM MM 2024 · 被引用 6 次
它引用的顶会 Paper5
- RNNLogic: Learning Logic Rules for Reasoning on Knowledge GraphsMeng Qu, Jun-Kun Chen, Louis-Pascal A. C. Xhonneux, Yoshua Bengio 等ICLR 2021 · 被引用 230 次
- Boosting Visual Question Answering with Context-aware Knowledge AggregationGuohao Li, Xin Wang, Wenwu ZhuACM MM 2020 · 被引用 82 次
- AVQA: A Dataset for Audio-Visual Question Answering on VideosPinci Yang, Xin Wang, Xuguang Duan, Hong Chen 等ACM MM 2022 · 被引用 60 次
- From Strings to Things: Knowledge-Enabled VQA Model That Can Read and ReasonAjeet Kumar Singh, Anand Mishra, Shashank Shekhar, Anirban ChakrabortyICCV 2019 · 被引用 54 次
- Hierarchical Conditional Relation Networks for Video Question AnsweringThao Minh Le, Vuong Le, Svetha Venkatesh, Truyen TranCVPR 2020
相关 Paper
- WFF: Wavelet-based Information Fusion for Multimodal Knowledge Graph Link PredictionXiaodi Xu, Lijie Li, Ye Wang, Tao Ren 等ACM MM 2025 · 被引用 1 次
- VL-KGE: Vision-Language Models Meet Knowledge Graph EmbeddingsAthanasios Efthymiou, Stevan Rudinac, Monika Kackovic, Nachoem Wijnberg 等WWW 2026 · 被引用 2 次
- Hybrid Transformer with Multi-level Fusion for Multimodal Knowledge Graph CompletionXiang Chen, Ningyu Zhang, Lei Li, Shumin Deng 等SIGIR 2022 · 被引用 227 次
- Multimodal Biological Knowledge Graph Completion via Triple Co-Attention MechanismDerong Xu, Jingbo Zhou, Tong Xu, Yuan Xia 等ICDE 2023 · 被引用 24 次
- InGram: Inductive Knowledge Graph Embedding via Relation GraphsJaejun Lee, Chanyoung Chung, Joyce Jiyoung WhangICML 2023 · 被引用 83 次
