Re-Attention for Visual Question Answering
Wenya Guo, Ying Zhang, Xiaoping Wu, Jufeng Yang, Xiangrui Cai, Xiaojie Yuan
Abstract
Visual Question Answering (VQA) requires a simultaneous understanding of images and questions. Existing methods achieve well performance by focusing on both key objects in images and key words in questions. However, the answer also contains rich information which can help to better describe the image and generate more accurate attention maps. In this paper, to utilize the information in answer, we propose a re-attention framework for the VQA task. We first associate image and question by calculating the similarity of each object-word pairs in the feature space. Then, based on the answer, the learned model re-attends the corresponding visual objects in images and reconstructs the initial attention map to produce consistent results. Benefiting from the re-attention procedure, the question can be better understood, and the satisfactory answer is generated. Extensive experiments on the benchmark dataset demonstrate the proposed method performs favorably against the state-of-the-art approaches.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a1c8af30-cd57-4b5b-a7d0-90e7d62e5e9eCited by top-tier papers5
- Focal and Composed Vision-semantic Modeling for Visual Question AnsweringYudong Han, Yangyang Guo, Jianhua Yin, Meng Liu et al.ACM MM 2021 · 14 citations
- Exploiting the Social-Like Prior in Transformer for Visual ReasoningYudong Han, Yupeng Hu, Xuemeng Song, Haoyu Tang et al.AAAI 2024 · 11 citations
- Scene Graph Generation Strategy with Co-occurrence Knowledge and Learnable Term FrequencyHyeongjin Kim, Sangwon Kim, Dasom Ahn, Jong Taek Lee et al.ICML 2024 · 8 citations
- Predicate Hierarchies Improve Few-Shot State ClassificationEmily Jin, Joy Hsu, Jiajun WuICLR 2025
- Prototype-Based Embedding Network for Scene Graph GenerationChaofan Zheng, Xinyu Lyu, Lianli Gao, Bo Dai et al.CVPR 2023
Related papers
- Multiple Objects-Aware Visual Question GenerationJiayuan Xie, Yi Cai, Qingbao Huang, Tao WangACM MM 2021 · 23 citations
- Multi-Question Learning for Visual Question AnsweringChenyi Lei, Lei Wu, Dong Liu, Zhao Li et al.AAAI 2020 · 9 citations
- REVIVE: Regional Visual Representation Matters in Knowledge-Based Visual Question AnsweringYuanze Lin, Yujia Xie, Dongdong Chen, Yichong Xu et al.NeurIPS 2022 · 119 citations
- Self-Bootstrapped Visual-Language Model for Knowledge Selection and Question AnsweringDongze Hao, Qunbo Wang, Longteng Guo, Jie Jiang et al.EMNLP 2024 · 4 citations
- Check It Again: Progressive Visual Question Answering via Visual EntailmentQingyi Si, Zheng Lin, Mingyu Zheng, Peng Fu et al.ACL 2021
