Improving Intra- and Inter-Modality Visual Relation for Image Captioning
Yong Wang, Wenkai Zhang, Qing Liu, Zhengyuan Zhang, Xin Gao, Xian Sun
Abstract
It is widely shared that capturing relationships among multi-modality features would be helpful for representing and ultimately describing an image. In this paper, we present a novel Intra- and Inter-modality visual Relation Transformer to improve connections among visual features, termed I2RT. Firstly, we propose Relation Enhanced Transformer Block (RETB) for image feature learning, which strengthens intra-modality visual relations among objects. Moreover, to bridge the gap between inter-modality feature representations, we align them explicitly via Visual Guided Alignment (VGA) module. Finally, an end-to-end formulation is adopted to train the whole model jointly. Experiments on the MS-COCO dataset show the effectiveness of our model, leading to improvements on all commonly used metrics on the "Karpathy" test split. Extensive ablation experiments are conducted for the comprehensive analysis of the proposed method.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 3cb78ef3-fa21-4276-9c1e-052c3f2ef1d2Related papers
- Multimodal Contrastive Training for Visual Representation LearningXin Yuan, Zhe Lin, Jason Kuen, Jianming Zhang et al.CVPR 2021
- M3TR: Multi-modal Multi-label Recognition with TransformerJiawei Zhao, Yifan Zhao, Jia LiACM MM 2021 · 45 citations
- Probing Inter-modality: Visual Parsing with Self-Attention for Vision-and-Language Pre-trainingHongwei Xue, Yupan Huang, Bei Liu, Houwen Peng et al.NeurIPS 2021 · 100 citations
- Modular Graph Transformer Networks for Multi-Label Image ClassificationHoang D. Nguyen, Xuan-Son Vu, Duc-Trong LeAAAI 2021 · 78 citations
- Transformer-based Dual Relation Graph for Multi-label Image RecognitionJiawei Zhao, Ke Yan, Yifan Zhao, Xiaowei Guo et al.ICCV 2021 · 109 citations
