Simple is not Easy: A Simple Strong Baseline for TextVQA and TextCaps
Qi Zhu, Chenyu Gao, Peng Wang, Qi Wu
摘要
Texts appearing in daily scenes that can be recognized by OCR (Optical Character Recognition) tools contain significant information, such as street name, product brand and prices. Two tasks -- text-based visual question answering and text-based image captioning, with a text extension from existing vision-language applications, are catching on rapidly. To address these problems, many sophisticated multi-modality encoding frameworks (such as heterogeneous graph structure) are being used. In this paper, we argue that a simple attention mechanism can do the same or even better job without any bells and whistles. Under this mechanism, we simply split OCR token features into separate visual- and linguistic-attention branches, and send them to a popular Transformer decoder to generate answers or captions. Surprisingly, we find this simple baseline model is rather strong -- it consistently outperforms state-of-the-art (SOTA) models on two popular benchmarks, TextVQA and all three tasks of ST-VQA, although these SOTA models use far more complex encoding mechanisms. Transferring it to text-based image captioning, we also surpass the TextCaps Challenge 2020 winner. We wish this work to set the new baseline for these two OCR text related applications and to inspire new thinking of multi-modality encoder design. Code is available at https://github.com/ZephyrZhuQi/ssbaseline
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- ViSTA: Vision and Scene Text Aggregation for Cross-Modal RetrievalMengjun Cheng, Yipeng Sun, Longchao Wang, Xiongwei Zhu 等CVPR 2022 · 被引用 86 次
- LaTr: Layout-Aware Transformer for Scene-Text VQAAli Furkan Biten, Ron Litman, Yusheng Xie, Srikar Appalaraju 等CVPR 2022 · 被引用 82 次
- Question-controlled Text-aware Image CaptioningAnwen Hu, Shizhe Chen, Qin JinACM MM 2021 · 被引用 11 次
- Locate Then Generate: Bridging Vision and Language with Bounding Box for Scene-Text VQAYongxin Zhu, Zhen Liu, Yukang Liang, Xin Li 等AAAI 2023 · 被引用 11 次
- Gather and Trace: Rethinking Video TextVQA from an Instance-oriented PerspectiveYan Zhang, Gangyan Zeng, Daiqing Wu, Huawen Shen 等ACM MM 2025 · 被引用 2 次
它引用的顶会 Paper3
- Scene Text Visual Question AnsweringAli Furkan Biten, Rubèn Tito, Andrés Mafla, Lluís Gómez i Bigorda 等ICCV 2019 · 被引用 482 次
- Multi-Modal Graph Neural Network for Joint Reasoning on Vision and Scene TextDifei Gao, Ke Li, Ruiping Wang, Shiguang Shan 等CVPR 2020
- Iterative Answer Prediction With Pointer-Augmented Multimodal Transformers for TextVQARonghang Hu, Amanpreet Singh, Trevor Darrell, Marcus RohrbachCVPR 2020
相关 Paper
- Beyond OCR + VQA: Involving OCR into the Flow for Robust and Accurate TextVQAGangyan Zeng, Yuan Zhang, Yu Zhou, Xiaomeng YangACM MM 2021 · 被引用 38 次
- TAP: Text-Aware Pre-Training for Text-VQA and Text-CaptionZhengyuan Yang, Yijuan Lu, Jianfeng Wang, Xi Yin 等CVPR 2021
- Separate and Locate: Rethink the Text in Text-based Visual Question AnsweringChengyang Fang, Jiangnan Li, Liang Li, Can Ma 等ACM MM 2023 · 被引用 18 次
- Towards Models that Can See and ReadRoy Ganz, Oren Nuriel, Aviad Aberdam, Yair Kittenplon 等ICCV 2023 · 被引用 17 次
- From Token to Word: OCR Token Evolution via Contrastive Learning and Semantic Matching for Text-VQAZan-Xia Jin, Mike Zheng Shou, Fang Zhou, Satoshi Tsutsui 等ACM MM 2022 · 被引用 11 次
