Iterative Back Modification for Faster Image Captioning
Zhengcong Fei
Abstract
Current state-of-the-art image captioning systems generally produce a sentence from left to right, and every step is conditioned on the given image and previously generated words. Nevertheless, such autoregressive nature makes the inference process difficult to parallelize and leads to high captioning latency. In this paper, we propose a non-autoregressive approach for faster image caption generation. Technically, low-dimension continuous latent variables are shaped to capture semantic information and word dependencies from extracted image features before sentence decoding. Moreover, we develop an iterative back modification inference algorithm, which continuously refines the latent variables with a look back mechanism and parallelly generates the whole sentence based on the updated latent variables in a constant number of steps. Extensive experiments demonstrate that our method achieves competitive performance compared to prevalent autoregressive captioning models while significantly reducing the decoding time on average.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 500ce7c8-46fa-4156-bb67-299a8ced33d0Cited by top-tier papers6
- Partially Non-Autoregressive Image CaptioningZhengcong FeiAAAI 2021 · 40 citations
- DeeCap: Dynamic Early Exiting for Efficient Image CaptioningZhengcong Fei, Xu Yan, Shuhui Wang, Qi TianCVPR 2022 · 39 citations
- Memory-Augmented Image CaptioningZhengcong FeiAAAI 2021 · 35 citations
- Semi-Autoregressive Image CaptioningXu Yan, Zhengcong Fei, Zekang Li, Shuhui Wang et al.ACM MM 2021 · 22 citations
- Uncertainty-Aware Image CaptioningZhengcong Fei, Mingyuan Fan, Li Zhu, Junshi Huang et al.AAAI 2023 · 21 citations
Related papers
- Latent-Variable Non-Autoregressive Neural Machine Translation with Deterministic Inference Using a Delta PosteriorRaphael Shu, Jason Lee, Hideki Nakayama, Kyunghyun ChoAAAI 2020 · 125 citations
- Efficient Modeling of Future Context for Image CaptioningZhengcong FeiACM MM 2022 · 10 citations
- Non-Autoregressive Coarse-to-Fine Video CaptioningBang Yang, Yuexian Zou, Fenglin Liu, Can ZhangAAAI 2021 · 92 citations
- Generating Image Descriptions via Sequential Cross-Modal Alignment Guided by Human GazeEce Takmaz, Sandro Pezzelle, Lisa Beinborn, Raquel FernándezEMNLP 2020 · 1 citation
- An EM Approach to Non-autoregressive Conditional Sequence GenerationZhiqing Sun, Yiming YangICML 2020 · 43 citations
