Ordered Attention for Coherent Visual Storytelling
Tom Braude, Idan Schwartz, Alexander G. Schwing, Ariel Shamir
摘要
We address the problem of visual storytelling, i.e., generating a story for a given sequence of images. While each story sentence should describe a corresponding image, a coherent story also needs to be consistent and relate to both future and past images. Current approaches encode images independently, disregarding relations between images. Our approach learns to encode images with different interactions based on the story position (i.e., past image or future image). To this end, we develop a novel message-passing-like algorithm for ordered image attention (OIA) that collects interactions across all the images in the sequence. Finally, to generate the story's sentences, a second attention mechanism picks the important image attention vectors with an Image-Sentence Attention (ISA). The obtained results improve the METEOR score on the VIST dataset by 1%. Furthermore, a thorough human study confirms improvements and demonstrates that order-based interactions significantly improve coherency (64.20% 28.70%). Source code available at ://github.com/tomateb/OIAVist.git
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Semantic Alignment for Multimodal Large Language ModelsTao Wu, Mengze Li, Jingyuan Chen, Wei Ji 等ACM MM 2024 · 被引用 13 次
- DataNarrative: Automated Data-Driven Storytelling with Visualizations and TextsMohammed Saidul Islam, Md. Tahmid Rahman Laskar, Md. Rizwan Parvez, Enamul Hoque 等EMNLP 2024 · 被引用 11 次
- Text-Only Training for Visual StorytellingYuechen Wang, Wengang Zhou, Zhenbo Lu, Houqiang LiACM MM 2023 · 被引用 4 次
- Multi-Modal Experience Inspired AI CreationQian Cao, Xu Chen, Ruihua Song, Hao Jiang 等ACM MM 2022 · 被引用 3 次
它引用的顶会 Paper9
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- Storytelling from an Image Stream Using Scene GraphsRuize Wang, Zhongyu Wei, Piji Li, Qi Zhang 等AAAI 2020 · 被引用 75 次
- Knowledge-Enriched Visual StorytellingChao-Chun Hsu, Zi-Yuan Chen, Chi-Yang Hsu, Chih-Chia Li 等AAAI 2020 · 被引用 53 次
- Latent Memory-augmented Graph Transformer for Visual StorytellingMengshi Qi, Jie Qin, Di Huang, Zhiqiang Shen 等ACM MM 2021 · 被引用 18 次
相关 Paper
- Imagine, Reason and Write: Visual Storytelling with Graph Knowledge and Relational ReasoningChunpu Xu, Min Yang, Chengming Li, Ying Shen 等AAAI 2021 · 被引用 39 次
- Hide-and-Tell: Learning to Bridge Photo Streams for Visual StorytellingYunjae Jung, Dahun Kim, Sanghyun Woo, Kyungsu Kim 等AAAI 2020 · 被引用 35 次
- ContextualStory: Consistent Visual Storytelling with Spatially-Enhanced and Storyline ContextSixiao Zheng, Yanwei FuAAAI 2025 · 被引用 12 次
- Intelligent Grimm - Open-ended Visual Storytelling via Latent Diffusion ModelsChang Liu, Haoning Wu, Yujie Zhong, Xiaoyun Zhang 等CVPR 2024 · 被引用 32 次
- A-CAP: Anticipation Captioning with Commonsense KnowledgeDuc Minh Vo, Quoc-An Luong, Akihiro Sugimoto, Hideki NakayamaCVPR 2023
