Towards Reasoning Ability in Scene Text Visual Question Answering
Qingqing Wang, Liqiang Xiao, Yue Lu, Yaohui Jin, Hao He
Abstract
Works on scene text visual question answering (TextVQA) always emphasize the importance of reasoning questions and image contents. However, we find current TextVQA models lack reasoning ability and tend to answer questions by exploiting dataset bias and language priors. Moreover, our observations indicate that recent accuracy improvement in TextVQA is mainly contributed by stronger OCR engines, better pre-training strategies and more Transformer layers, instead of newly proposed networks. In this work, towards the reasoning ability, we 1) conduct module-wise contribution analysis to quantitatively investigate how existing works improve accuracies in TextVQA; 2) design a gradient-based explainability method to explore why TextVQA models answer what they answer and find evidence for their predictions; 3) perform qualitative experiments to visually analyze models reasoning ability and explore potential reasons behind such a poor ability.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers3
- LaTr: Layout-Aware Transformer for Scene-Text VQAAli Furkan Biten, Ron Litman, Yusheng Xie, Srikar Appalaraju et al.CVPR 2022 · 82 citations
- Towards Models that Can See and ReadRoy Ganz, Oren Nuriel, Aviad Aberdam, Yair Kittenplon et al.ICCV 2023 · 17 citations
- AdaCoder: Adaptive Prompt Compression for Programmatic Visual Question AnsweringMahiro Ukai, Shuhei Kurita, Atsushi Hashimoto, Yoshitaka Ushiku et al.ACM MM 2024 · 1 citation
Related papers
- Beyond OCR + VQA: Involving OCR into the Flow for Robust and Accurate TextVQAGangyan Zeng, Yuan Zhang, Yu Zhou, Xiaomeng YangACM MM 2021 · 38 citations
- From Strings to Things: Knowledge-Enabled VQA Model That Can Read and ReasonAjeet Kumar Singh, Anand Mishra, Shashank Shekhar, Anirban ChakrabortyICCV 2019 · 54 citations
- Track the Answer: Extending TextVQA from Image to Video with Spatio-Temporal CluesYan Zhang, Gangyan Zeng, Huawen Shen, Daiqing Wu et al.AAAI 2025 · 1 citation
- Do Vision-Language Transformers Exhibit Visual Commonsense? An Empirical Study of VCRZhenyang Li, Yangyang Guo, Kejie Wang, Xiaolin Chen et al.ACM MM 2023 · 11 citations
- Cascade Reasoning Network for Text-based Visual Question AnsweringFen Liu, Guanghui Xu, Qi Wu, Qing Du et al.ACM MM 2020 · 61 citations
