Latent Memory-augmented Graph Transformer for Visual Storytelling
Mengshi Qi, Jie Qin, Di Huang, Zhiqiang Shen, Yi Yang, Jiebo Luo
Abstract
Visual storytelling aims to automatically generate a human-like short story given an image stream. Most existing works utilize either scene-level or object-level representations, neglecting the interaction among objects in each image and the sequential dependency between consecutive images. In this paper, we present a novel Latent Memory-augmented Graph Transformer (LMGT ), a Transformer based framework for visual story generation. LMGT directly inherits the merits from the Transformer, which is further enhanced with two carefully designed components, i.e., a graph encoding module and a latent memory unit. Specifically, the graph encoding module exploits the semantic relationships among image regions and attentively aggregates critical visual features based on the parsed scene graphs. Furthermore, to better preserve inter-sentence coherence and topic consistency, we introduce an augmented latent memory unit that learns and records highly summarized latent information as the story line from the image stream and the sentence history. Experimental results on three widely-used datasets demonstrate the superior performance of LMGT over the state-of-the-art methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d2c5dee6-c753-47bc-93fd-5e3a1f3bf464Cited by top-tier papers5
- With a Little Help from your own Past: Prototypical Memory Networks for Image CaptioningManuele Barraco, Sara Sarto, Marcella Cornia, Lorenzo Baraldi et al.ICCV 2023 · 33 citations
- Ordered Attention for Coherent Visual StorytellingTom Braude, Idan Schwartz, Alexander G. Schwing, Ariel ShamirACM MM 2022 · 12 citations
- Towards Balanced Multi-Modal Learning in 3D Human Pose EstimationMengshi Qi, Jiaxuan Peng, Xianlin Zhang, Huadong MaCVPR 2026 · 12 citations
- Robo-SGG: Exploiting Layout-Oriented Normalization and Restitution Can Improve Robust Scene Graph GenerationChangsheng Lv, Zijian Fu, Mengshi QiCVPR 2026 · 4 citations
- Text-Only Training for Visual StorytellingYuechen Wang, Wengang Zhou, Zhenbo Lu, Houqiang LiACM MM 2023 · 4 citations
Builds on17
- VideoBERT: A Joint Model for Video and Language Representation LearningChen Sun, Austin Myers, Carl Vondrick, Kevin Murphy et al.ICCV 2019 · 1,396 citations
- Attention on Attention for Image CaptioningLun Huang, Wenmin Wang, Jie Chen, Xiaoyong WeiICCV 2019 · 992 citations
- Entangled Transformer for Image CaptioningGuang Li, Linchao Zhu, Ping Liu, Yi YangICCV 2019 · 346 citations
- MART: Memory-Augmented Recurrent Transformer for Coherent Video Paragraph CaptioningJie Lei, Liwei Wang, Yelong Shen, Dong Yu et al.ACL 2020 · 168 citations
- Storytelling from an Image Stream Using Scene GraphsRuize Wang, Zhongyu Wei, Piji Li, Qi Zhang et al.AAAI 2020 · 75 citations
Related papers
- MMT: Image-guided Story Ending Generation with Multimodal Memory TransformerDizhan Xue, Shengsheng Qian, Quan Fang, Changsheng XuACM MM 2022 · 16 citations
- Story Visualization by Online Text Augmentation with Context MemoryDaechul Ahn, Daneul Kim, Gwangmo Song, Seung Hwan Kim et al.ICCV 2023 · 11 citations
- LayoutTransformer: Scene Layout Generation With Conceptual and Spatial DiversityCheng-Fu Yang, Wan-Cyuan Fan, Fu-En Yang, Yu-Chiang Frank WangCVPR 2021
- Imagine, Reason and Write: Visual Storytelling with Graph Knowledge and Relational ReasoningChunpu Xu, Min Yang, Chengming Li, Ying Shen et al.AAAI 2021 · 39 citations
- ContextualStory: Consistent Visual Storytelling with Spatially-Enhanced and Storyline ContextSixiao Zheng, Yanwei FuAAAI 2025 · 12 citations
