Improving OCR-Based Image Captioning by Incorporating Geometrical Relationship
Jing Wang, Jinhui Tang, Mingkun Yang, Xiang Bai, Jiebo Luo
摘要
OCR-based image captioning aims to automatically describe images based on all the visual entities (both visual objects and scene text) in images. Compared with conventional image captioning, the reasoning of scene text is required for OCR-based image captioning since the generated descriptions often contain multiple OCR tokens. Existing methods attempt to achieve this goal via encoding the OCR tokens with rich visual and semantic representations. However, strong correlations between OCR tokens may not be established with such limited representations. In this paper, we propose to enhance the connections between OCR tokens from the viewpoint of exploiting the geometrical relationship. We comprehensively consider the height, width, distance, IoU and orientation relations between the OCR tokens for constructing the geometrical relationship. To integrate the learned relation as well as the visual and semantic representations into a unified framework, a Long Short-Term Memory plus Relation-aware pointer network (LSTM-R) architecture is presented in this paper. Under the guidance of the geometrical relationship between OCR tokens, our LSTM-R capitalizes on a newly-devised relationaware pointer network to select OCR tokens from the scene text for OCR-based image captioning. Extensive experiments demonstrate the effectiveness of our LSTM-R. More remarkably, LSTM-R achieves state-of-the-art performance on TextCaps, with the CIDEr-D score being increased from 98.0% to 109.3%.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper6
- Attention on Attention for Image CaptioningLun Huang, Wenmin Wang, Jie Chen, Xiaoyong WeiICCV 2019 · 被引用 992 次
- Entangled Transformer for Image CaptioningGuang Li, Linchao Zhu, Ping Liu, Yi YangICCV 2019 · 被引用 346 次
- Hierarchy Parsing for Image CaptioningTing Yao, Yingwei Pan, Yehao Li, Tao MeiICCV 2019 · 被引用 183 次
- Multimodal Attention with Image Text Spatial Relationship for OCR-Based Image CaptioningJing Wang, Jinhui Tang, Jiebo LuoACM MM 2020 · 被引用 55 次
- Meshed-Memory Transformer for Image CaptioningMarcella Cornia, Matteo Stefanini, Lorenzo Baraldi, Rita CucchiaraCVPR 2020
相关 Paper
- Direction Relation Transformer for Image CaptioningZeliang Song, Xiaofei Zhou, Linhua Dong, Jianlong Tan 等ACM MM 2021 · 被引用 31 次
- TAP: Text-Aware Pre-Training for Text-VQA and Text-CaptionZhengyuan Yang, Yijuan Lu, Jianfeng Wang, Xi Yin 等CVPR 2021
- Confidence-aware Non-repetitive Multimodal Transformers for TextCapsZhaokai Wang, Renda Bao, Qi Wu, Si LiuAAAI 2021 · 被引用 28 次
- Zero-TextCap: Zero-shot Framework for Text-based Image CaptioningDongsheng Xu, Wenye Zhao, Yi Cai, Qingbao HuangACM MM 2023 · 被引用 4 次
- Boosting Weakly Supervised Referring Image Segmentation via Progressive ComprehensionZaiquan Yang, Yuhao Liu, Jiaying Lin, Gerhard P. Hancke 等NeurIPS 2024 · 被引用 14 次
