Memory-Augmented Image Captioning
Zhengcong Fei
摘要
Current deep learning-based image captioning systems have been proven to store practical knowledge with their parameters and achieve competitive performances in the public datasets. Nevertheless, their ability to access and precisely manipulate the mastered knowledge is still limited. Besides, providing evidence for decisions and updating memory information are also important yet under explored. Towards this goal, we introduce a memory-augmented method, which extends an existing image caption model by incorporating extra explicit knowledge from a memory bank. Adequate knowledge is recalled according to the similarity distance in the embedding space of history context, and the memory bank can be constructed conveniently from any matched image-text set, e.g., the previous training data. Incorporating such non-parametric memory-augmented method to various captioning baselines, the performance of resulting captioners imporves consistently on the evaluation benchmark. More encouragingly, extensive experiments demonstrate that our approach holds the capability for efficiently adapting to larger training datasets, by simply transferring the memory bank without any additional training.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Implicit Identity Representation Conditioned Memory Compensation Network for Talking Head Video GenerationFa-Ting Hong, Dan XuICCV 2023 · 被引用 75 次
- Attention-Aligned Transformer for Image CaptioningZhengcong FeiAAAI 2022 · 被引用 42 次
- DeeCap: Dynamic Early Exiting for Efficient Image CaptioningZhengcong Fei, Xu Yan, Shuhui Wang, Qi TianCVPR 2022 · 被引用 39 次
- Accelerating Retrieval-Augmented GenerationDerrick Quinn, Mohammad Nouri, Neel Patel, John Salihu 等ASPLOS 2025 · 被引用 37 次
- Efficient Modeling of Future Context for Image CaptioningZhengcong FeiACM MM 2022 · 被引用 10 次
它引用的顶会 Paper5
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Attention on Attention for Image CaptioningLun Huang, Wenmin Wang, Jie Chen, Xiaoyong WeiICCV 2019 · 被引用 992 次
- Show, Recall, and Tell: Image Captioning with Recall MechanismLi Wang, Zechen Bai, Yonghua Zhang, Hongtao LuAAAI 2020 · 被引用 73 次
- Iterative Back Modification for Faster Image CaptioningZhengcong FeiACM MM 2020 · 被引用 27 次
- Meshed-Memory Transformer for Image CaptioningMarcella Cornia, Matteo Stefanini, Lorenzo Baraldi, Rita CucchiaraCVPR 2020
相关 Paper
- MemCap: Memorizing Style Knowledge for Image CaptioningWentian Zhao, Xinxiao Wu, Xiaoxun ZhangAAAI 2020 · 被引用 86 次
- Evcap: Retrieval-Augmented Image Captioning with External Visual-Name Memory for Open-World ComprehensionJiaxuan Li, Duc Minh Vo, Akihiro Sugimoto, Hideki NakayamaCVPR 2024
- MeaCap: Memory-Augmented Zero-shot Image CaptioningZequn Zeng, Yan Xie, Hao Zhang, Chiyu Chen 等CVPR 2024 · 被引用 38 次
- MuRAG: Multimodal Retrieval-Augmented Generator for Open Question Answering over Images and TextWenhu Chen, Hexiang Hu, Xi Chen, Pat Verga 等EMNLP 2022 · 被引用 89 次
- Towards flexible perception with visual memoryRobert Geirhos, Priyank Jaini, Austin Stone, Sourabh Medapati 等ICML 2025
