Towards Accurate Text-Based Image Captioning With Content Diversity Exploration
Guanghui Xu, Shuaicheng Niu, Mingkui Tan, Yucheng Luo, Qing Du, Qi Wu
Abstract
Text-based image captioning (TextCap) which aims to read and reason images with texts is crucial for a machine to understand a detailed and complex scene environment, considering that texts are omnipresent in daily life. This task, however, is very challenging because an image often contains complex texts and visual information that is hard to be described comprehensively. Existing methods attempt to extend the traditional image captioning methods to solve this task, which focus on describing the overall scene of images by one global caption. This is infeasible because the complex text and visual information cannot be described well within one caption. To resolve this difficulty, we seek to generate multiple captions that accurately describe different parts of an image in detail. To achieve this purpose, there are three key challenges: 1) it is hard to decide which parts of the texts of images to copy or paraphrase; 2) it is non-trivial to capture the complex relationship between diverse texts in an image; 3) how to generate multiple captions with diverse content is still an open problem. To conquer these, we propose a novel Anchor-Captioner method. Specifically, we first find the important tokens which are supposed to be paid more attention to and consider them as anchors. Then, for each chosen anchor, we group its relevant texts to construct the corresponding anchor-centred graph (ACG). Last, based on different ACGs, we conduct the multi-view caption generation to improve the content diversity of generated captions. Experimental results show that our method not only achieves SOTA performance but also generates diverse captions to describe images.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext abd9b330-e9f1-44cb-a1bb-63903134e89cCited by top-tier papers8
- Efficient Test-Time Model Adaptation without ForgettingShuaicheng Niu, Jiaxiang Wu, Yifan Zhang, Yaofo Chen et al.ICML 2022 · 579 citations
- Towards End-to-End Unified Scene Text Detection and Layout AnalysisShangbang Long, Siyang Qin, Dmitry Panteleev, Alessandro Bissacco et al.CVPR 2022 · 86 citations
- MAGIC: Multimodal relAtional Graph adversarIal inferenCe for Diverse and Unpaired Text-Based Image CaptioningWenqiao Zhang, Haochen Shi, Jiannan Guo, Shengyu Zhang et al.AAAI 2022 · 52 citations
- GSRFormer: Grounded Situation Recognition Transformer with Alternate Semantic Attention RefinementZhi-Qi Cheng, Qi Dai, Siyao Li, Teruko Mitamura et al.ACM MM 2022 · 40 citations
- DeeCap: Dynamic Early Exiting for Efficient Image CaptioningZhengcong Fei, Xu Yan, Shuhui Wang, Qi TianCVPR 2022 · 39 citations
Builds on8
- Attention on Attention for Image CaptioningLun Huang, Wenmin Wang, Jie Chen, Xiaoyong WeiICCV 2019 · 992 citations
- Scene Text Visual Question AnsweringAli Furkan Biten, Rubèn Tito, Andrés Mafla, Lluís Gómez i Bigorda et al.ICCV 2019 · 482 citations
- Location-Aware Graph Convolutional Networks for Video Question AnsweringDeng Huang, Peihao Chen, Runhao Zeng, Qing Du et al.AAAI 2020 · 187 citations
- Multimodal Attention with Image Text Spatial Relationship for OCR-Based Image CaptioningJing Wang, Jinhui Tang, Jiebo LuoACM MM 2020 · 55 citations
- How to Train Your Agent to Read and WriteLi Liu, Mengge He, Guanghui Xu, Mingkui Tan et al.AAAI 2021 · 3 citations
Related papers
- CapOnImage: Context-driven Dense-Captioning on ImageYiqi Gao, Xinglin Hou, Yuanmeng Zhang, Tiezheng Ge et al.EMNLP 2022 · 5 citations
- Zero-TextCap: Zero-shot Framework for Text-based Image CaptioningDongsheng Xu, Wenye Zhao, Yi Cai, Qingbao HuangACM MM 2023 · 4 citations
- Set Prediction Guided by Semantic Concepts for Diverse Video CaptioningYifan Lu, Ziqi Zhang, Chunfeng Yuan, Peng Li et al.AAAI 2024 · 7 citations
- Question-controlled Text-aware Image CaptioningAnwen Hu, Shizhe Chen, Qin JinACM MM 2021 · 11 citations
- Open-Book Video Captioning With Retrieve-Copy-Generate NetworkZiqi Zhang, Zhongang Qi, Chunfeng Yuan, Ying Shan et al.CVPR 2021
