Commonsense Knowledge Aware Concept Selection For Diverse and Informative Visual Storytelling
Hong Chen, Yifei Huang, Hiroya Takamura, Hideki Nakayama
Abstract
Visual storytelling is a task of generating relevant and interesting stories for given image sequences. In this work we aim at increasing the diversity of the generated stories while preserving the informative content from the images. We propose to foster the diversity and informativeness of a generated story by using a concept selection module that suggests a set of concept candidates. Then, we utilize a large scale pre-trained model to convert concepts and images into full stories. To enrich the candidate concepts, a commonsense knowledge graph is created for each image sequence from which the concept candidates are proposed. To obtain appropriate concepts from the graph, we propose two novel modules that consider the correlation among candidate concepts and the image-concept correlation. Extensive automatic and human evaluation results demonstrate that our model can produce reasonable concepts. This enables our model to outperform the previous models by a large margin on the diversity and informativeness of the story, while retaining the relevance of the story to the image sequence.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Detecting and Grounding Important Characters in Visual StoriesDanyang Liu, Frank KellerAAAI 2023 · 11 citations
- Leveraging Weak Cross-Modal Guidance for Coherence Modelling via Iterative LearningYi Bin, Junrong Liao, Yujuan Ding, Haoxuan Li et al.ACM MM 2024 · 3 citations
- Weakly Supervised Temporal Sentence Grounding with Uncertainty-Guided Self-trainingYifei Huang, Lijin Yang, Yoichi SatoCVPR 2023
- A-CAP: Anticipation Captioning with Commonsense KnowledgeDuc Minh Vo, Quoc-An Luong, Akihiro Sugimoto, Hideki NakayamaCVPR 2023
Builds on6
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- Story Realization: Expanding Plot Events into SentencesPrithviraj Ammanabrolu, Ethan Tien, Wesley Cheung, Zhaochen Luo et al.AAAI 2020 · 79 citations
- Storytelling from an Image Stream Using Scene GraphsRuize Wang, Zhongyu Wei, Piji Li, Qi Zhang et al.AAAI 2020 · 75 citations
- What Makes A Good Story? Designing Composite Rewards for Visual StorytellingJunjie Hu, Yu Cheng, Zhe Gan, Jingjing Liu et al.AAAI 2020 · 73 citations
- Knowledge-Enriched Visual StorytellingChao-Chun Hsu, Zi-Yuan Chen, Chi-Yang Hsu, Chih-Chia Li et al.AAAI 2020 · 53 citations
Related papers
- Imagine, Reason and Write: Visual Storytelling with Graph Knowledge and Relational ReasoningChunpu Xu, Min Yang, Chengming Li, Ying Shen et al.AAAI 2021 · 39 citations
- KM-BART: Knowledge Enhanced Multimodal BART for Visual Commonsense GenerationYiran Xing, Zai Shi, Zhao Meng, Gerhard Lakemeyer et al.ACL 2021
- Retrieve, Caption, Generate: Visual Grounding for Enhancing Commonsense in Text Generation ModelsSteven Y. Feng, Kevin Lu, Zhuofu Tao, Malihe Alikhani et al.AAAI 2022 · 15 citations
- Generated Knowledge Prompting for Commonsense ReasoningJiacheng Liu, Alisa Liu, Ximing Lu, Sean Welleck et al.ACL 2022
- Text-Only Training for Visual StorytellingYuechen Wang, Wengang Zhou, Zhenbo Lu, Houqiang LiACM MM 2023 · 4 citations
