Hierarchy Parsing for Image Captioning
Ting Yao, Yingwei Pan, Yehao Li, Tao Mei
摘要
It is always well believed that parsing an image into constituent visual patterns would be helpful for understanding and representing an image. Nevertheless, there has not been evidence in support of the idea on describing an image with a natural-language utterance. In this paper, we introduce a new design to model a hierarchy from instance level (segmentation), region level (detection) to the whole image to delve into a thorough image understanding for captioning. Specifically, we present a HIerarchy Parsing (HIP) architecture that novelly integrates hierarchical structure into image encoder. Technically, an image decomposes into a set of regions and some of the regions are resolved into finer ones. Each region then regresses to an instance, i.e., foreground of the region. Such process naturally builds a hierarchal tree. A tree-structured Long Short-Term Memory (Tree-LSTM) network is then employed to interpret the hierarchal structure and enhance all the instance-level, region-level and image-level features. Our HIP is appealing in view that it is pluggable to any neural captioning models. Extensive experiments on COCO image captioning dataset demonstrate the superiority of HIP. More remarkably, HIP plus a top-down attention-based LSTM decoder increases CIDEr-D performance from 120.1% to 127.2% on COCO Karpathy test split. When further endowing instance-level and region-level features from HIP with semantic relation learnt through Graph Convolutional Networks (GCN), CIDEr-D is boosted up to 130.6%.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper25
- Dual-level Collaborative Transformer for Image CaptioningYunpeng Luo, Jiayi Ji, Xiaoshuai Sun, Liujuan Cao 等AAAI 2021 · 被引用 349 次
- VisualGPT: Data-efficient Adaptation of Pretrained Language Models for Image CaptioningJun Chen, Han Guo, Kai Yi, Boyang Li 等CVPR 2022 · 被引用 169 次
- Injecting Semantic Concepts into End-to-End Image CaptioningZhiyuan Fang, Jianfeng Wang, Xiaowei Hu, Lin Liang 等CVPR 2022 · 被引用 125 次
- Comprehending and Ordering Semantics for Image CaptioningYehao Li, Yingwei Pan, Ting Yao, Tao MeiCVPR 2022 · 被引用 124 次
- Beyond a Pre-Trained Object Detector: Cross-Modal Textual and Visual Context for Image CaptioningChia-Wen Kuo, Zsolt KiraCVPR 2022 · 被引用 88 次
相关 Paper
- Progressive Tree-Structured Prototype Network for End-to-End Image CaptioningPengpeng Zeng, Jinkuan Zhu, Jingkuan Song, Lianli GaoACM MM 2022 · 被引用 35 次
- HAAV: Hierarchical Aggregation of Augmented Views for Image CaptioningChia-Wen Kuo, Zsolt KiraCVPR 2023
- Direction Relation Transformer for Image CaptioningZeliang Song, Xiaofei Zhou, Linhua Dong, Jianlong Tan 等ACM MM 2021 · 被引用 31 次
- Exploring Overall Contextual Information for Image Captioning in Human-Like Cognitive StyleHongwei Ge, Zehang Yan, Kai Zhang, Mingde Zhao 等ICCV 2019 · 被引用 25 次
- Improving OCR-Based Image Captioning by Incorporating Geometrical RelationshipJing Wang, Jinhui Tang, Mingkun Yang, Xiang Bai 等CVPR 2021
