A-CAP: Anticipation Captioning with Commonsense Knowledge
Duc Minh Vo, Quoc-An Luong, Akihiro Sugimoto, Hideki Nakayama
摘要
Humans possess the capacity to reason about the future based on a sparse collection of visual cues acquired over time. In order to emulate this ability, we introduce a novel task called Anticipation Captioning, which generates a caption for an unseen oracle image using a sparsely temporally-ordered set of images. To tackle this new task, we propose a model called A-CAP, which incorporates commonsense knowledge into a pre-trained vision-language model, allowing it to anticipate the caption. Through both qualitative and quantitative evaluations on a customized visual storytelling dataset, A-CAP outperforms other image captioning methods and establishes a strong baseline for anticipation captioning. We also address the challenges inherent in this task.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper11
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Conditional Prompt Learning for Vision-Language ModelsKaiyang Zhou, Jingkang Yang, Chen Change Loy, Ziwei LiuCVPR 2022 · 被引用 1,438 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- CLIPScore: A Reference-free Evaluation Metric for Image CaptioningJack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras 等EMNLP 2021 · 被引用 937 次
相关 Paper
- Imagine, Reason and Write: Visual Storytelling with Graph Knowledge and Relational ReasoningChunpu Xu, Min Yang, Chengming Li, Ying Shen 等AAAI 2021 · 被引用 39 次
- Joint Commonsense and Relation Reasoning for Image and Video CaptioningJingyi Hou, Xinxiao Wu, Xiaoxun Zhang, Yayun Qi 等AAAI 2020 · 被引用 52 次
- Commonsense Knowledge Aware Concept Selection For Diverse and Informative Visual StorytellingHong Chen, Yifei Huang, Hiroya Takamura, Hideki NakayamaAAAI 2021 · 被引用 49 次
- Hide-and-Tell: Learning to Bridge Photo Streams for Visual StorytellingYunjae Jung, Dahun Kim, Sanghyun Woo, Kyungsu Kim 等AAAI 2020 · 被引用 35 次
- Unicoder-VL: A Universal Encoder for Vision and Language by Cross-Modal Pre-TrainingGen Li, Nan Duan, Yuejian Fang, Ming Gong 等AAAI 2020 · 被引用 966 次
