Distilled Cross-Combination Transformer for Image Captioning with Dual Refined Visual Features
Junbo Hu, Zhixin Li
摘要
Transformer-based encoders that encode both region and grid features are the preferred choice for the image captioning task due to their multi-head self-attention mechanism. This mechanism ensures superior capture of relationships and contextual information between various regions in an image. However, because of the Transformer block stacking, self-attention computes the visual features several times, increasing computing costs and producing a great deal of redundant feature calculation. In this paper, we propose a novel Distilled Cross-Combination Transformer (DCCT) network. Specifically, we first design a distillation cascade fusion encoder(DCFE) to filter out redundant features in visual features that affect attentional focus, obtaining refined features. Additionally, we introduce a parallel cross-fusion attention module (PCFA) that fully utilizes the complementarity and correlation between grid and region features to better fuse the encoded dual visual features. Extensive experiments on the MSCOCO dataset demonstrate that the proposed DCCT strategy outperforms many state-of-the-art techniques and attains exceptional performance.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper1
问问它们各自怎么用它相关 Paper
- Dual-level Collaborative Transformer for Image CaptioningYunpeng Luo, Jiayi Ji, Xiaoshuai Sun, Liujuan Cao 等AAAI 2021 · 被引用 349 次
- DIFNet: Boosting Visual Information Flow for Image CaptioningMingrui Wu, Xuying Zhang, Xiaoshuai Sun, Yiyi Zhou 等CVPR 2022 · 被引用 68 次
- Meshed-Memory Transformer for Image CaptioningMarcella Cornia, Matteo Stefanini, Lorenzo Baraldi, Rita CucchiaraCVPR 2020
- CropCap: Embedding Visual Cross-Partition Dependency for Image CaptioningBo Wang, Zhao Zhang, Suiyi Zhao, Haijun Zhang 等ACM MM 2023 · 被引用 9 次
- End-to-End Transformer Based Model for Image CaptioningYiyu Wang, Jungang Xu, Yingfei SunAAAI 2022 · 被引用 178 次
