Prophet Attention: Predicting Attention with Future Attention
Fenglin Liu, Xuancheng Ren, Xian Wu, Shen Ge, Wei Fan, Yuexian Zou, Xu Sun
Abstract
Recently, attention based models have been used extensively in many sequence-tosequence learning systems. Especially for image captioning, the attention based models are expected to ground correct image regions with proper generated words. However, for each time step in the decoding process, the attention based models usually use the hidden state of the current input to attend to the image regions. Under this setting, these attention models have a "deviated focus" problem that they calculate the attention weights based on previous words instead of the one to be generated, impairing the performance of both grounding and captioning. In this paper, we propose the Prophet Attention, similar to the form of self-supervision. In the training stage, this module utilizes the future information to calculate the "ideal" attention weights towards image regions. These calculated "ideal" weights are further used to regularize the "deviated" attention. In this manner, image regions are grounded with the correct words. The proposed Prophet Attention can be easily incorporated into existing image captioning models to improve their performance of both grounding and captioning. The experiments on the Flickr30k Entities and the MSCOCO datasets show that the proposed Prophet Attention consistently outperforms baselines in both automatic metrics and human evaluations. It is worth noticing that we set new state-of-the-arts on the two benchmark datasets and achieve the 1st place on the leaderboard of the online MSCOCO benchmark in terms of the default ranking score, i.e., CIDEr-c40.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 26ea0d72-bc22-4436-9391-b18d7e1dc873Cited by top-tier papers7
- Expectation-Maximization Contrastive Learning for Compact Video-and-Language RepresentationsPeng Jin, Jinfa Huang, Fenglin Liu, Xian Wu et al.NeurIPS 2022 · 105 citations
- With a Little Help from your own Past: Prototypical Memory Networks for Image CaptioningManuele Barraco, Sara Sarto, Marcella Cornia, Lorenzo Baraldi et al.ICCV 2023 · 33 citations
- Distributed Attention for Grounded Image CaptioningNenglun Chen, Xingjia Pan, Runnan Chen, Lei Yang et al.ACM MM 2021 · 19 citations
- MultiCapCLIP: Auto-Encoding Prompts for Zero-Shot Multilingual Visual CaptioningBang Yang, Fenglin Liu, Xian Wu, Yaowei Wang et al.ACL 2023 · 10 citations
- Weakly-Supervised Generation and Grounding of Visual Descriptions with Conditional Generative ModelsEffrosyni Mavroudi, René VidalCVPR 2022 · 5 citations
Builds on10
- PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive SummarizationJingqing Zhang, Yao Zhao, Mohammad Saleh, Peter J. LiuICML 2020 · 2,453 citations
- Attention on Attention for Image CaptioningLun Huang, Wenmin Wang, Jie Chen, Xiaoyong WeiICCV 2019 · 992 citations
- Taking a HINT: Leveraging Explanations to Make Vision and Language Models More GroundedRamprasaath Ramasamy Selvaraju, Stefan Lee, Yilin Shen, Hongxia Jin et al.ICCV 2019 · 288 citations
- Auto-Encoding Knowledge Graph for Unsupervised Medical Report GenerationFenglin Liu, Chenyu You, Xian Wu, Shen Ge et al.NeurIPS 2021 · 144 citations
- Federated Learning for Vision-and-Language Grounding ProblemsFenglin Liu, Xian Wu, Shen Ge, Wei Fan et al.AAAI 2020 · 130 citations
Related papers
- Attention-Aligned Transformer for Image CaptioningZhengcong FeiAAAI 2022 · 42 citations
- Image Captioning with Context-Aware Auxiliary GuidanceZeliang Song, Xiaofei Zhou, Zhendong Mao, Jianlong TanAAAI 2021 · 36 citations
- Discrete-Continuous Action Space Policy Gradient-Based Attention for Image-Text MatchingShiyang Yan, Li Yu, Yuan XieCVPR 2021
- Trust Prophet or Not? Taking a Further Verification Step toward Accurate Scene Text RecognitionAnna Zhu, Ke Xiao, Bo Zhou, Runmin WangACM MM 2024 · 6 citations
- More Grounded Image Captioning by Distilling Image-Text Matching ModelYuanen Zhou, Meng Wang, Daqing Liu, Zhenzhen Hu et al.CVPR 2020
