Distilled Cross-Combination Transformer for Image Captioning with Dual Refined Visual Features
Junbo Hu, Zhixin Li
Abstract
Transformer-based encoders that encode both region and grid features are the preferred choice for the image captioning task due to their multi-head self-attention mechanism. This mechanism ensures superior capture of relationships and contextual information between various regions in an image. However, because of the Transformer block stacking, self-attention computes the visual features several times, increasing computing costs and producing a great deal of redundant feature calculation. In this paper, we propose a novel Distilled Cross-Combination Transformer (DCCT) network. Specifically, we first design a distillation cascade fusion encoder(DCFE) to filter out redundant features in visual features that affect attentional focus, obtaining refined features. Additionally, we introduce a parallel cross-fusion attention module (PCFA) that fully utilizes the complementarity and correlation between grid and region features to better fuse the encoded dual visual features. Extensive experiments on the MSCOCO dataset demonstrate that the proposed DCCT strategy outperforms many state-of-the-art techniques and attains exceptional performance.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 4d398138-cec9-4b92-b6c5-9d255edceb95Cited by top-tier papers1
Ask how each one uses itRelated papers
- Dual-level Collaborative Transformer for Image CaptioningYunpeng Luo, Jiayi Ji, Xiaoshuai Sun, Liujuan Cao et al.AAAI 2021 · 349 citations
- DIFNet: Boosting Visual Information Flow for Image CaptioningMingrui Wu, Xuying Zhang, Xiaoshuai Sun, Yiyi Zhou et al.CVPR 2022 · 68 citations
- Meshed-Memory Transformer for Image CaptioningMarcella Cornia, Matteo Stefanini, Lorenzo Baraldi, Rita CucchiaraCVPR 2020
- CropCap: Embedding Visual Cross-Partition Dependency for Image CaptioningBo Wang, Zhao Zhang, Suiyi Zhao, Haijun Zhang et al.ACM MM 2023 · 9 citations
- End-to-End Transformer Based Model for Image CaptioningYiyu Wang, Jungang Xu, Yingfei SunAAAI 2022 · 178 citations
