CapOnImage: Context-driven Dense-Captioning on Image
Yiqi Gao, Xinglin Hou, Yuanmeng Zhang, Tiezheng Ge, Yuning Jiang, Peng Wang
摘要
Existing image captioning systems are dedicated to generating narrative captions for images, which are spatially detached from theimage in presentation. However, texts can also be used as decorations on the image to highlight the key points and increase theattractiveness of images. In this work, we introduce a new taskcalled captioning on image (CapOnImage), which aims to generatedense captions at different locations of the image based on contextual information. To fully exploit the surrounding visual context togenerate the most suitable caption for each location, we propose amulti-modal pre-training model with multi-level pre-training tasksthat progressively learn the correspondence between texts and image locations from easy to difficult. Since the model may generateredundant captions for nearby locations, we further enhance thelocation embedding with neighbor locations as context. For thisnew task, we also introduce a large-scale benchmark called CapOnImage2M, which contains 2.1 million product images, each with anaverage of 4.8 spatially localized captions. Compared with other image captioning model variants, our model achieves the best resultsin both captioning accuracy and diversity aspects.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- AutoPoster: A Highly Automatic and Content-aware Design System for Advertising Poster GenerationJinpeng Lin, Min Zhou, Ye Ma, Yifan Gao 等ACM MM 2023 · 被引用 24 次
- TextPainter: Multimodal Text Image Generation with Visual-harmony and Text-comprehension for Poster DesignYifan Gao, Jinpeng Lin, Min Zhou, Chuanbin Liu 等ACM MM 2023 · 被引用 6 次
- CapDet: Unifying Dense Captioning and Open-World Detection PretrainingYanxin Long, Youpeng Wen, Jianhua Han, Hang Xu 等CVPR 2023
它引用的顶会 Paper11
- Attention on Attention for Image CaptioningLun Huang, Wenmin Wang, Jie Chen, Xiaoyong WeiICCV 2019 · 被引用 992 次
- Unicoder-VL: A Universal Encoder for Vision and Language by Cross-Modal Pre-TrainingGen Li, Nan Duan, Yuejian Fang, Ming Gong 等AAAI 2020 · 被引用 966 次
- End-to-End Dense Video Captioning with Parallel DecodingTeng Wang, Ruimao Zhang, Zhichao Lu, Feng Zheng 等ICCV 2021 · 被引用 238 次
- LayoutVAE: Stochastic Scene Layout Generation From a Label SetAkash Abdu Jyothi, Thibaut Durand, Jiawei He, Leonid Sigal 等ICCV 2019 · 被引用 194 次
- TAP: Text-Aware Pre-Training for Text-VQA and Text-CaptionZhengyuan Yang, Yijuan Lu, Jianfeng Wang, Xi Yin 等CVPR 2021
相关 Paper
- LocCa: Visual Pretraining with Location-aware CaptionersBo Wan, Michael Tschannen, Yongqin Xian, Filip Pavetic 等NeurIPS 2024 · 被引用 38 次
- Towards Accurate Text-Based Image Captioning With Content Diversity ExplorationGuanghui Xu, Shuaicheng Niu, Mingkui Tan, Yucheng Luo 等CVPR 2021
- MAGIC: Multimodal relAtional Graph adversarIal inferenCe for Diverse and Unpaired Text-Based Image CaptioningWenqiao Zhang, Haochen Shi, Jiannan Guo, Shengyu Zhang 等AAAI 2022 · 被引用 52 次
- Pixel Aligned Language ModelsJiarui Xu, Xingyi Zhou, Shen Yan, Xiuye Gu 等CVPR 2024 · 被引用 6 次
- Multimodal Attention with Image Text Spatial Relationship for OCR-Based Image CaptioningJing Wang, Jinhui Tang, Jiebo LuoACM MM 2020 · 被引用 55 次
