In Defense of Grid Features for Visual Question Answering
Huaizu Jiang, Ishan Misra, Marcus Rohrbach, Erik G. Learned-Miller, Xinlei Chen
摘要
Popularized as 'bottom-up' attention [2], bounding box (or region) based visual features have recently surpassed vanilla grid-based convolutional features as the de facto standard for vision and language tasks like visual question answering (VQA). However, it is not clear whether the advantages of regions (e.g. better localization) are the key reasons for the success of bottom-up attention. In this paper, we revisit grid features for VQA, and find they can work surprisingly well -running more than an order of magnitude faster with the same accuracy (e.g. if pre-trained in a similar fashion). Through extensive experiments, we verify that this observation holds true across different VQA models (reporting a state-of-the-art accuracy on VQA 2.0 test-std, 72.71), datasets, and generalizes well to other tasks like image captioning. As grid features make the model design and training process much simpler, this enables us to train them end-to-end and also use a more flexible network design. We learn VQA models end-to-end, from pixels directly to answers, and show that strong performance is achievable without using any region annotations in pre-training. We hope our findings help further improve the scientific understanding and the practical application of VQA. Code and features will be made available. * This work was done when Huaizu Jiang was an intern at FAIR. 1 We use the terms 'region' and 'bounding box' interchangeably.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper67
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 被引用 2,258 次
- How Much Can CLIP Benefit Vision-and-Language Tasks?Sheng Shen, Liunian Harold Li, Hao Tan, Mohit Bansal 等ICLR 2022 · 被引用 503 次
- FLAVA: A Foundational Language And Vision Alignment ModelAmanpreet Singh, Ronghang Hu, Vedanuj Goswami, Guillaume Couairon 等CVPR 2022 · 被引用 483 次
- MERLOT: Multimodal Neural Script Knowledge ModelsRowan Zellers, Ximing Lu, Jack Hessel, Youngjae Yu 等NeurIPS 2021 · 被引用 463 次
- Mesh GraphormerKevin Lin, Lijuan Wang, Zicheng LiuICCV 2021 · 被引用 399 次
它引用的顶会 Paper2
相关 Paper
- Simple is not Easy: A Simple Strong Baseline for TextVQA and TextCapsQi Zhu, Chenyu Gao, Peng Wang, Qi WuAAAI 2021 · 被引用 59 次
- Visual Commonsense R-CNNTan Wang, Jianqiang Huang, Hanwang Zhang, Qianru SunCVPR 2020
- REVIVE: Regional Visual Representation Matters in Knowledge-Based Visual Question AnsweringYuanze Lin, Yujia Xie, Dongdong Chen, Yichong Xu 等NeurIPS 2022 · 被引用 119 次
- RSTNet: Captioning With Adaptive Attention on Visual and Non-Visual WordsXuying Zhang, Xiaoshuai Sun, Yunpeng Luo, Jiayi Ji 等CVPR 2021
- Sentence Attention Blocks for Answer GroundingSeyedalireza Khoshsirat, Chandra KambhamettuICCV 2023 · 被引用 8 次
