MMT: Image-guided Story Ending Generation with Multimodal Memory Transformer
Dizhan Xue, Shengsheng Qian, Quan Fang, Changsheng Xu
Abstract
As a specific form of story generation, Image-guided Story Ending Generation (IgSEG) is a recently proposed task of generating a story ending for a given multi-sentence story plot and an ending-related image. Unlike existing image captioning tasks or story ending generation tasks, IgSEG aims to generate a factual description that conforms to both the contextual logic and the relevant visual concepts. To date, existing methods for IgSEG ignore the relationships between the multimodal information and do not integrate multimodal features appropriately. Therefore, in this work, we propose Multimodal Memory Transformer (MMT), an end-to-end framework that models and fuses both contextual and visual information to effectively capture the multimodal dependency for IgSEG. Firstly, we extract textual and visual features separately by employing modality-specific large-scale pretrained encoders. Secondly, we utilize the memory-augmented cross-modal attention network to learn cross-modal relationships and conduct the fine-grained feature fusion effectively. Finally, a multimodal transformer decoder constructs attention among multimodal features to learn the story dependency and generates informative, reasonable, and coherent story endings. In experiments, extensive automatic evaluation results and human evaluation results indicate the significant performance boost of our proposed MMT over state-of-the-art methods on two benchmark datasets.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 44e902ca-fcb5-45ab-a9df-84547e7ed44eCited by top-tier papers4
- With a Little Help from your own Past: Prototypical Memory Networks for Image CaptioningManuele Barraco, Sara Sarto, Marcella Cornia, Lorenzo Baraldi et al.ICCV 2023 · 33 citations
- Variational Causal Inference Network for Explanatory Visual Question AnsweringDizhan Xue, Shengsheng Qian, Changsheng XuICCV 2023 · 19 citations
- Semantic Alignment for Multimodal Large Language ModelsTao Wu, Mengze Li, Jingyuan Chen, Wei Ji et al.ACM MM 2024 · 13 citations
- BadAgent: Inserting and Activating Backdoor Attacks in LLM AgentsYifei Wang, Dizhan Xue, Shengjie Zhang, Shengsheng QianACL 2024
Related papers
- Latent Memory-augmented Graph Transformer for Visual StorytellingMengshi Qi, Jie Qin, Di Huang, Zhiqiang Shen et al.ACM MM 2021 · 18 citations
- MMTN: Multi-Modal Memory Transformer Network for Image-Report Consistent Medical Report GenerationYiming Cao, Lizhen Cui, Lei Zhang, Fuqiang Yu et al.AAAI 2023 · 56 citations
- Imagine, Reason and Write: Visual Storytelling with Graph Knowledge and Relational ReasoningChunpu Xu, Min Yang, Chengming Li, Ying Shen et al.AAAI 2021 · 39 citations
- Cross-Counter-Repeat Attention for Enhanced Understanding of Visual Semantics in Radiology Report GenerationXiaolei Bo, Feiyang Yang, Feilong Xu, Xiaoli ZhangACM MM 2025
- Story Visualization by Online Text Augmentation with Context MemoryDaechul Ahn, Daneul Kim, Gwangmo Song, Seung Hwan Kim et al.ICCV 2023 · 11 citations
