In Defense of Grid Features for Visual Question Answering
Huaizu Jiang, Ishan Misra, Marcus Rohrbach, Erik G. Learned-Miller, Xinlei Chen
Abstract
Popularized as 'bottom-up' attention [2], bounding box (or region) based visual features have recently surpassed vanilla grid-based convolutional features as the de facto standard for vision and language tasks like visual question answering (VQA). However, it is not clear whether the advantages of regions (e.g. better localization) are the key reasons for the success of bottom-up attention. In this paper, we revisit grid features for VQA, and find they can work surprisingly well -running more than an order of magnitude faster with the same accuracy (e.g. if pre-trained in a similar fashion). Through extensive experiments, we verify that this observation holds true across different VQA models (reporting a state-of-the-art accuracy on VQA 2.0 test-std, 72.71), datasets, and generalizes well to other tasks like image captioning. As grid features make the model design and training process much simpler, this enables us to train them end-to-end and also use a more flexible network design. We learn VQA models end-to-end, from pixels directly to answers, and show that strong performance is achievable without using any region annotations in pre-training. We hope our findings help further improve the scientific understanding and the practical application of VQA. Code and features will be made available. * This work was done when Huaizu Jiang was an intern at FAIR. 1 We use the terms 'region' and 'bounding box' interchangeably.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7190d1df-04e6-4d57-8974-e50fe0669ca2Cited by top-tier papers67
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 2,258 citations
- How Much Can CLIP Benefit Vision-and-Language Tasks?Sheng Shen, Liunian Harold Li, Hao Tan, Mohit Bansal et al.ICLR 2022 · 503 citations
- FLAVA: A Foundational Language And Vision Alignment ModelAmanpreet Singh, Ronghang Hu, Vedanuj Goswami, Guillaume Couairon et al.CVPR 2022 · 483 citations
- MERLOT: Multimodal Neural Script Knowledge ModelsRowan Zellers, Ximing Lu, Jack Hessel, Youngjae Yu et al.NeurIPS 2021 · 463 citations
- Mesh GraphormerKevin Lin, Lijuan Wang, Zicheng LiuICCV 2021 · 399 citations
Builds on2
Related papers
- Simple is not Easy: A Simple Strong Baseline for TextVQA and TextCapsQi Zhu, Chenyu Gao, Peng Wang, Qi WuAAAI 2021 · 59 citations
- Visual Commonsense R-CNNTan Wang, Jianqiang Huang, Hanwang Zhang, Qianru SunCVPR 2020
- REVIVE: Regional Visual Representation Matters in Knowledge-Based Visual Question AnsweringYuanze Lin, Yujia Xie, Dongdong Chen, Yichong Xu et al.NeurIPS 2022 · 119 citations
- RSTNet: Captioning With Adaptive Attention on Visual and Non-Visual WordsXuying Zhang, Xiaoshuai Sun, Yunpeng Luo, Jiayi Ji et al.CVPR 2021
- Sentence Attention Blocks for Answer GroundingSeyedalireza Khoshsirat, Chandra KambhamettuICCV 2023 · 8 citations
