Iterative Back Modification for Faster Image Captioning
Zhengcong Fei
摘要
Current state-of-the-art image captioning systems generally produce a sentence from left to right, and every step is conditioned on the given image and previously generated words. Nevertheless, such autoregressive nature makes the inference process difficult to parallelize and leads to high captioning latency. In this paper, we propose a non-autoregressive approach for faster image caption generation. Technically, low-dimension continuous latent variables are shaped to capture semantic information and word dependencies from extracted image features before sentence decoding. Moreover, we develop an iterative back modification inference algorithm, which continuously refines the latent variables with a look back mechanism and parallelly generates the whole sentence based on the updated latent variables in a constant number of steps. Extensive experiments demonstrate that our method achieves competitive performance compared to prevalent autoregressive captioning models while significantly reducing the decoding time on average.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper6
- Partially Non-Autoregressive Image CaptioningZhengcong FeiAAAI 2021 · 被引用 40 次
- DeeCap: Dynamic Early Exiting for Efficient Image CaptioningZhengcong Fei, Xu Yan, Shuhui Wang, Qi TianCVPR 2022 · 被引用 39 次
- Memory-Augmented Image CaptioningZhengcong FeiAAAI 2021 · 被引用 35 次
- Semi-Autoregressive Image CaptioningXu Yan, Zhengcong Fei, Zekang Li, Shuhui Wang 等ACM MM 2021 · 被引用 22 次
- Uncertainty-Aware Image CaptioningZhengcong Fei, Mingyuan Fan, Li Zhu, Junshi Huang 等AAAI 2023 · 被引用 21 次
相关 Paper
- Latent-Variable Non-Autoregressive Neural Machine Translation with Deterministic Inference Using a Delta PosteriorRaphael Shu, Jason Lee, Hideki Nakayama, Kyunghyun ChoAAAI 2020 · 被引用 125 次
- Efficient Modeling of Future Context for Image CaptioningZhengcong FeiACM MM 2022 · 被引用 10 次
- Non-Autoregressive Coarse-to-Fine Video CaptioningBang Yang, Yuexian Zou, Fenglin Liu, Can ZhangAAAI 2021 · 被引用 92 次
- Generating Image Descriptions via Sequential Cross-Modal Alignment Guided by Human GazeEce Takmaz, Sandro Pezzelle, Lisa Beinborn, Raquel FernándezEMNLP 2020 · 被引用 1 次
- An EM Approach to Non-autoregressive Conditional Sequence GenerationZhiqing Sun, Yiming YangICML 2020 · 被引用 43 次
