Non-autoregressive Sequence-to-Sequence Vision-Language Models
Kunyu Shi, Qi Dong, Luis Goncalves, Zhuowen Tu, Stefano Soatto
Abstract
Sequence-to-sequence vision-language models are showing promise, but their applicability is limited by their inference latency due to their autoregressive way of generating predictions. We propose a parallel decoding sequence-to-sequence vision-language model, trained with a Query-CTC loss, that marginalizes over multiple inference paths in the decoder. This allows us to model the joint distribution of tokens, rather than restricting to conditional distribution as in an autoregressive model. The resulting model, NARVL, achieves performance on-par with its state-of-the-art autoregressive counterpart, but is faster at inference time, reducing from the linear complexity associated with the sequential generation of tokens to a paradigm of constant time joint inference.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d7ea3d97-eb4c-4aa7-b05f-c88129653e68Cited by top-tier papers2
- Ask and Remember: A Questions-Only Replay Strategy for Continual Visual Question AnsweringImad Eddine Marouf, Enzo Tartaglione, Stéphane Lathuilière, Joost van de WeijerICCV 2025 · 4 citations
- Enhancing Vision-Language Pre-Training with Rich SupervisionsYuan Gao, Kunyu Shi, Pengkai Zhu, Edouard Belval et al.CVPR 2024
Builds on17
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- MDETR - Modulated Detection for End-to-End Multi-Modal UnderstandingAishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve et al.ICCV 2021 · 1,114 citations
- OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning FrameworkPeng Wang, An Yang, Rui Men, Junyang Lin et al.ICML 2022 · 1,058 citations
- Unifying Vision-and-Language Tasks via Text GenerationJaemin Cho, Jie Lei, Hao Tan, Mohit BansalICML 2021 · 624 citations
- Large-Scale Adversarial Training for Vision-and-Language Representation LearningZhe Gan, Yen-Chun Chen, Linjie Li, Chen Zhu et al.NeurIPS 2020 · 561 citations
Related papers
- An EM Approach to Non-autoregressive Conditional Sequence GenerationZhiqing Sun, Yiming YangICML 2020 · 43 citations
- MAGVLT: Masked Generative Vision-and-Language TransformerSungwoong Kim, Daejin Jo, Donghoon Lee, Jongmin KimCVPR 2023
- DREAM: Drafting with Refined Target Features and Entropy-Adaptive Cross-Attention Fusion for Multimodal Speculative DecodingYunhai Hu, Tianhua Xia, Zining Liu, Rahul Raman et al.NeurIPS 2025 · 15 citations
- Parallelized Autoregressive Visual GenerationYuqing Wang, Shuhuai Ren, Zhijie Lin, Yujin Han et al.CVPR 2025
- Improving Non-Autoregressive Translation Models Without DistillationXiao Shi Huang, Felipe Pérez, Maksims VolkovsICLR 2022 · 60 citations
