Lune

ACM MM2020Top-tier venue

Improving Intra- and Inter-Modality Visual Relation for Image Captioning

Yong Wang, Wenkai Zhang, Qing Liu, Zhengyuan Zhang, Xin Gao, Xian Sun

2020Year
24Citations

Abstract

It is widely shared that capturing relationships among multi-modality features would be helpful for representing and ultimately describing an image. In this paper, we present a novel Intra- and Inter-modality visual Relation Transformer to improve connections among visual features, termed I2RT. Firstly, we propose Relation Enhanced Transformer Block (RETB) for image feature learning, which strengthens intra-modality visual relations among objects. Moreover, to bridge the gap between inter-modality feature representations, we align them explicitly via Visual Guided Alignment (VGA) module. Finally, an end-to-end formulation is adopted to train the whole model jointly. Experiments on the MS-COCO dataset show the effectiveness of our model, leading to improvements on all commonly used metrics on the "Karpathy" test split. Extensive ablation experiments are conducted for the comprehensive analysis of the proposed method.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

lune papers get 3cb78ef3-fa21-4276-9c1e-052c3f2ef1d2

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines