DeeCap: Dynamic Early Exiting for Efficient Image Captioning
Zhengcong Fei, Xu Yan, Shuhui Wang, Qi Tian
摘要
Both accuracy and efficiency are crucial for image captioning in real-world scenarios. Although Transformer-based models have gained significant improved captioning performance, their computational cost is very high. A feasible way to reduce the time complexity is to exit the prediction early in internal decoding layers without passing the entire model. However, it is not straightforward to devise early exiting into image captioning due to the following issues. On one hand, the representation in shallow layers lacks high-level semantic and sufficient cross-modal fusion information for accurate prediction. On the other hand, the exiting decisions made by internal classifiers are unreliable sometimes. To solve these issues, we propose DeeCap framework for efficient image captioning, which dynamically selects proper-sized decoding layers from a global perspective to exit early. The key to successful early exiting lies in the specially designed imitation learning mechanism, which predicts the deep layer activation with shallow layer features. By deliberately merging the imitation learning into the whole image captioning architecture, the imitated deep layer representation can mitigate the loss brought by the missing of actual deep layers when early exiting is undertaken, resulting in significant reduction in calculation cost with small sacrifice of accuracy. Experiments on the MS COCO and Flickr30k datasets demonstrate the DeeCap can achieve competitive performances with 4× speed-up. Code is available at: https://github.com/feizc/DeeCap.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- DeeR-VLA: Dynamic Inference of Multimodal Large Language Models for Efficient Robot ExecutionYang Yue, Yulin Wang, Bingyi Kang, Yizeng Han 等NeurIPS 2024 · 被引用 153 次
- Fast yet Safe: Early-Exiting with Risk ControlMetod Jazbec, Alexander Timans, Tin Hadzi Veljkovic, Kaspar Sakmann 等NeurIPS 2024 · 被引用 35 次
- Uncertainty-Aware Image CaptioningZhengcong Fei, Mingyuan Fan, Li Zhu, Junshi Huang 等AAAI 2023 · 被引用 21 次
- Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical FindingsQiong Wu, Wenhao Lin, Yiyi Zhou, Weihao Ye 等NeurIPS 2025 · 被引用 16 次
- Alternating Updates for Efficient TransformersCenk Baykal, Dylan J. Cutler, Nishanth Dikkala, Nikhil Ghosh 等NeurIPS 2023 · 被引用 13 次
它引用的顶会 Paper22
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Unified Vision-Language Pre-Training for Image Captioning and VQALuowei Zhou, Hamid Palangi, Lei Zhang, Houdong Hu 等AAAI 2020 · 被引用 1,047 次
- Attention on Attention for Image CaptioningLun Huang, Wenmin Wang, Jie Chen, Xiaoyong WeiICCV 2019 · 被引用 992 次
- BERT Loses Patience: Fast and Robust Inference with Early ExitWangchunshu Zhou, Canwen Xu, Tao Ge, Julian J. McAuley 等NeurIPS 2020 · 被引用 473 次
- Dual-level Collaborative Transformer for Image CaptioningYunpeng Luo, Jiayi Ji, Xiaoshuai Sun, Liujuan Cao 等AAAI 2021 · 被引用 349 次
相关 Paper
- You Need Multiple Exiting: Dynamic Early Exiting for Accelerating Unified Vision Language ModelShengkun Tang, Yaqing Wang, Zhenglun Kong, Tianchi Zhang 等CVPR 2023
- Reject Decoding via Language-Vision Models for Text-to-Image SynthesisFuxiang Wu, Liu Liu, Fusheng Hao, Fengxiang He 等AAAI 2023 · 被引用 2 次
- Improving Image Captioning by Leveraging Intra- and Inter-layer Global Representation in Transformer NetworkJiayi Ji, Yunpeng Luo, Xiaoshuai Sun, Fuhai Chen 等AAAI 2021 · 被引用 206 次
- Semi-Autoregressive Image CaptioningXu Yan, Zhengcong Fei, Zekang Li, Shuhui Wang 等ACM MM 2021 · 被引用 22 次
- Accurate and Fast Compressed Video CaptioningYaojie Shen, Xin Gu, Kai Xu, Heng Fan 等ICCV 2023 · 被引用 53 次
