KM-BART: Knowledge Enhanced Multimodal BART for Visual Commonsense Generation
Yiran Xing, Zai Shi, Zhao Meng, Gerhard Lakemeyer, Yunpu Ma, Roger Wattenhofer
Abstract
We present Knowledge Enhanced Multimodal BART (KM-BART), which is a Transformerbased sequence-to-sequence model capable of reasoning about commonsense knowledge from multimodal inputs of images and texts. We adapt the generative BART architecture (Lewis et al., 2020) to a multimodal model with visual and textual inputs. We further develop novel pretraining tasks to improve the model performance on the Visual Commonsense Generation (VCG) task. In particular, our pretraining task of Knowledge-based Commonsense Generation (KCG) boosts model performance on the VCG task by leveraging commonsense knowledge from a large language model pretrained on external commonsense knowledge graphs. To the best of our knowledge, we are the first to propose a dedicated task for improving model performance on the VCG task. Experimental results show that our model reaches state-of-the-art performance on the VCG task (Park et al., 2020) by applying these novel pretraining tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0987025d-890d-4c19-ad49-740fd2849e69Cited by top-tier papers9
- Vision-Language Pre-Training for Multimodal Aspect-Based Sentiment AnalysisYan Ling, Jianfei Yu, Rui XiaACL 2022 · 116 citations
- Improving Cross-Modal Alignment with Synthetic Pairs for Text-Only Image CaptioningZhiyue Liu, Jinyuan Liu, Fanrong MaAAAI 2024 · 23 citations
- UniSA: Unified Generative Framework for Sentiment AnalysisZaijing Li, Ting-En Lin, Yuchuan Wu, Meng Liu et al.ACM MM 2023 · 22 citations
- Summary-Oriented Vision Modeling for Multimodal Abstractive SummarizationYunlong Liang, Fandong Meng, Jinan Xu, Jiaan Wang et al.ACL 2023 · 17 citations
- Multi-source Semantic Graph-based Multimodal Sarcasm Explanation GenerationLiqiang Jing, Xuemeng Song, Kun Ouyang, Mengzhao Jia et al.ACL 2023 · 17 citations
Builds on3
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- Unified Vision-Language Pre-Training for Image Captioning and VQALuowei Zhou, Hamid Palangi, Lei Zhang, Houdong Hu et al.AAAI 2020 · 1,047 citations
- Visual Semantic Reasoning for Image-Text MatchingKunpeng Li, Yulun Zhang, Kai Li, Yuanyuan Li et al.ICCV 2019 · 598 citations
Related papers
- Retrieve, Caption, Generate: Visual Grounding for Enhancing Commonsense in Text Generation ModelsSteven Y. Feng, Kevin Lu, Zhuofu Tao, Malihe Alikhani et al.AAAI 2022 · 15 citations
- KG-BART: Knowledge Graph-Augmented BART for Generative Commonsense ReasoningYe Liu, Yao Wan, Lifang He, Hao Peng et al.AAAI 2021 · 220 citations
- SGEITL: Scene Graph Enhanced Image-Text Learning for Visual Commonsense ReasoningZhecan Wang, Haoxuan You, Liunian Harold Li, Alireza Zareian et al.AAAI 2022 · 40 citations
- Multiple Knowledge Syncretic Transformer for Natural Dialogue GenerationXiangyu Zhao, Longbiao Wang, Ruifang He, Ting Yang et al.WWW 2020 · 27 citations
- Simple Yet Effective: Structure Guided Pre-trained Transformer for Multi-modal Knowledge Graph ReasoningKe Liang, Lingyuan Meng, Yue Liu, Meng Liu et al.ACM MM 2024 · 29 citations
