Towards Reasoning Ability in Scene Text Visual Question Answering
Qingqing Wang, Liqiang Xiao, Yue Lu, Yaohui Jin, Hao He
摘要
Works on scene text visual question answering (TextVQA) always emphasize the importance of reasoning questions and image contents. However, we find current TextVQA models lack reasoning ability and tend to answer questions by exploiting dataset bias and language priors. Moreover, our observations indicate that recent accuracy improvement in TextVQA is mainly contributed by stronger OCR engines, better pre-training strategies and more Transformer layers, instead of newly proposed networks. In this work, towards the reasoning ability, we 1) conduct module-wise contribution analysis to quantitatively investigate how existing works improve accuracies in TextVQA; 2) design a gradient-based explainability method to explore why TextVQA models answer what they answer and find evidence for their predictions; 3) perform qualitative experiments to visually analyze models reasoning ability and explore potential reasons behind such a poor ability.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper3
- LaTr: Layout-Aware Transformer for Scene-Text VQAAli Furkan Biten, Ron Litman, Yusheng Xie, Srikar Appalaraju 等CVPR 2022 · 被引用 82 次
- Towards Models that Can See and ReadRoy Ganz, Oren Nuriel, Aviad Aberdam, Yair Kittenplon 等ICCV 2023 · 被引用 17 次
- AdaCoder: Adaptive Prompt Compression for Programmatic Visual Question AnsweringMahiro Ukai, Shuhei Kurita, Atsushi Hashimoto, Yoshitaka Ushiku 等ACM MM 2024 · 被引用 1 次
相关 Paper
- Beyond OCR + VQA: Involving OCR into the Flow for Robust and Accurate TextVQAGangyan Zeng, Yuan Zhang, Yu Zhou, Xiaomeng YangACM MM 2021 · 被引用 38 次
- From Strings to Things: Knowledge-Enabled VQA Model That Can Read and ReasonAjeet Kumar Singh, Anand Mishra, Shashank Shekhar, Anirban ChakrabortyICCV 2019 · 被引用 54 次
- Track the Answer: Extending TextVQA from Image to Video with Spatio-Temporal CluesYan Zhang, Gangyan Zeng, Huawen Shen, Daiqing Wu 等AAAI 2025 · 被引用 1 次
- Do Vision-Language Transformers Exhibit Visual Commonsense? An Empirical Study of VCRZhenyang Li, Yangyang Guo, Kejie Wang, Xiaolin Chen 等ACM MM 2023 · 被引用 11 次
- Cascade Reasoning Network for Text-based Visual Question AnsweringFen Liu, Guanghui Xu, Qi Wu, Qing Du 等ACM MM 2020 · 被引用 61 次
