Multimodal Attention with Image Text Spatial Relationship for OCR-Based Image Captioning
Jing Wang, Jinhui Tang, Jiebo Luo
Abstract
OCR-based image captioning is the task of automatically describing images based on reading and understanding written text contained in images. Compared to conventional image captioning, this task is more challenging, especially when the image contains multiple text tokens and visual objects. The difficulties originate from how to make full use of the knowledge contained in the textual entities to facilitate sentence generation and how to predict a text token based on the limited information provided by the image. Such problems are not yet fully investigated in existing research. In this paper, we present a novel design - Multimodal Attention Captioner with OCR Spatial Relationship (dubbed as MMA-SR) architecture, which manages information from different modalities with a multimodal attention network and explores spatial relationships between text tokens for OCR-based image captioning. Specifically, the representations of text tokens and objects are fed into a three-layer LSTM captioner. Different attention scores for text tokens and objects are exploited through the multimodal attention network. Based on the attended features and the LSTM states, words are selected from the common vocabulary or from the image text by incorporating the learned spatial relationships between text tokens. Extensive experiments conducted on the TextCaps dataset verify the effectiveness of the proposed MMA-SR method. More remarkably, our MMA-SR increases CIDEr-D score from 93.7% to 98.0%.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers10
- Pix2Struct: Screenshot Parsing as Pretraining for Visual Language UnderstandingKenton Lee, Mandar Joshi, Iulia Raluca Turc, Hexiang Hu et al.ICML 2023 · 426 citations
- Controllable Image Captioning via PromptingNing Wang, Jiahao Xie, Jihao Wu, Mingbo Jia et al.AAAI 2023 · 43 citations
- NICGSlowDown: Evaluating the Efficiency Robustness of Neural Image Caption Generation ModelsSimin Chen, Zihe Song, Mirazul Haque, Cong Liu et al.CVPR 2022 · 34 citations
- Latent Memory-augmented Graph Transformer for Visual StorytellingMengshi Qi, Jie Qin, Di Huang, Zhiqiang Shen et al.ACM MM 2021 · 18 citations
- Towards Models that Can See and ReadRoy Ganz, Oren Nuriel, Aviad Aberdam, Yair Kittenplon et al.ICCV 2023 · 17 citations
Related papers
- Improving OCR-Based Image Captioning by Incorporating Geometrical RelationshipJing Wang, Jinhui Tang, Mingkun Yang, Xiang Bai et al.CVPR 2021
- Confidence-aware Non-repetitive Multimodal Transformers for TextCapsZhaokai Wang, Renda Bao, Qi Wu, Si LiuAAAI 2021 · 28 citations
- TAP: Text-Aware Pre-Training for Text-VQA and Text-CaptionZhengyuan Yang, Yijuan Lu, Jianfeng Wang, Xi Yin et al.CVPR 2021
- Attention on Attention for Image CaptioningLun Huang, Wenmin Wang, Jie Chen, Xiaoyong WeiICCV 2019 · 992 citations
- Direction Relation Transformer for Image CaptioningZeliang Song, Xiaofei Zhou, Linhua Dong, Jianlong Tan et al.ACM MM 2021 · 31 citations
