Imagine, Reason and Write: Visual Storytelling with Graph Knowledge and Relational Reasoning
Chunpu Xu, Min Yang, Chengming Li, Ying Shen, Xiang Ao, Ruifeng Xu
Abstract
Visual storytelling is the task of generating a short story to describe an ordered image stream. Different from visual captions, stories contain not only factual descriptions, but also imaginary concepts that do not appear in the images. In this paper, we propose a novel imagine-reason-write generation framework (IRW) for visual storytelling, inspired by the logic of humans when they write a story. First, a multimodal imagining module is leveraged to learn the imaginative storyline explicitly, improving the coherence and reasonability of the generated story. Second, we employ a relational reasoning module to fully exploit the external knowledge (commonsense knowledge base) and task-specific knowledge (scene graph and event graph) with a relational reasoning method based on the storyline. In this way, we can effectively capture the most informative commonsense and visual relationships among objects in images, enhancing the diversity and informativeness of the generated story. Finally, we integrate the visual information and semantic (concept) information to generate human-like stories. Extensive experiments on a benchmark dataset (i.e., VIST) demonstrate that the proposed IRW framework substantially outperforms the state-of-the-art methods across multiple evaluation metrics.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bc13aa17-b267-4975-8093-8ee1593d7640Cited by top-tier papers1
Ask how each one uses itBuilds on3
- Storytelling from an Image Stream Using Scene GraphsRuize Wang, Zhongyu Wei, Piji Li, Qi Zhang et al.AAAI 2020 · 75 citations
- What Makes A Good Story? Designing Composite Rewards for Visual StorytellingJunjie Hu, Yu Cheng, Zhe Gan, Jingjing Liu et al.AAAI 2020 · 73 citations
- Hide-and-Tell: Learning to Bridge Photo Streams for Visual StorytellingYunjae Jung, Dahun Kim, Sanghyun Woo, Kyungsu Kim et al.AAAI 2020 · 35 citations
Related papers
- Commonsense Knowledge Aware Concept Selection For Diverse and Informative Visual StorytellingHong Chen, Yifei Huang, Hiroya Takamura, Hideki NakayamaAAAI 2021 · 49 citations
- A-CAP: Anticipation Captioning with Commonsense KnowledgeDuc Minh Vo, Quoc-An Luong, Akihiro Sugimoto, Hideki NakayamaCVPR 2023
- Retrieve, Caption, Generate: Visual Grounding for Enhancing Commonsense in Text Generation ModelsSteven Y. Feng, Kevin Lu, Zhuofu Tao, Malihe Alikhani et al.AAAI 2022 · 15 citations
- LogiStory: A Logic-Aware Framework for Multi-Image Story VisualizationChutian Meng, Fan Ma, Chi Zhang, Jiaxu Miao et al.ICLR 2026 · 3 citations
- Text-Only Training for Visual StorytellingYuechen Wang, Wengang Zhou, Zhenbo Lu, Houqiang LiACM MM 2023 · 4 citations
