LANTERN: Accelerating Visual Autoregressive Models with Relaxed Speculative Decoding
Doohyuk Jang, Sihwan Park, June Yong Yang, Yeonsung Jung, Jihun Yun, Souvik Kundu, Sungyub Kim, Eunho Yang
Abstract
Auto-Regressive (AR) models have recently gained prominence in image generation, often matching or even surpassing the performance of diffusion models. However, one major limitation of AR models is their sequential nature, which processes tokens one at a time, slowing down generation compared to models like GANs or diffusion-based methods that operate more efficiently. While speculative decoding has proven effective for accelerating LLMs by generating multiple tokens in a single forward, its application in visual AR models remains largely unexplored. In this work, we identify a challenge in this setting, which we term token selection ambiguity, wherein visual AR models frequently assign uniformly low probabilities to tokens, hampering the performance of speculative decoding. To overcome this challenge, we propose a relaxed acceptance condition referred to as LANTERN that leverages the interchangeability of tokens in latent space. This relaxation restores the effectiveness of speculative decoding in visual AR models by enabling more flexible use of candidate tokens that would otherwise be prematurely rejected. Furthermore, by incorporating a total variation distance bound, we ensure that these speed gains are achieved without significantly compromising image quality or semantic coherence. Experimental results demonstrate the efficacy of our method in providing a substantial speed-up over speculative decoding. In specific, compared to a naïve application of the state-of-the-art speculative decoding, LANTERN increases speed-ups by 1.75× and 1.82×, as compared to greedy decoding and random sampling, respectively, when applied to LlamaGen, a contemporary visual AR model. The code is publicly available at https://github.com/jadohu/LANTERN .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bf10f776-399f-433f-b2aa-8e4c940b7e78Cited by top-tier papers22
- Towards Better & Faster Autoregressive Image Generation: From the Perspective of EntropyXiaoxiao Ma, Feng Zhao, Pengyang Ling, Haibo Qiu et al.NeurIPS 2025 · 12 citations
- Speculative Jacobi-Denoising Decoding for Accelerating Autoregressive Text-to-image GenerationYao Teng, Fuyun Wang, Xian Liu, Zhekai Chen et al.NeurIPS 2025 · 8 citations
- FreqExit: Enabling Early-Exit Inference for Visual Autoregressive Models via Frequency-Aware GuidanceYing Li, Chengfei Lyu, Huan WangNeurIPS 2025 · 6 citations
- VVS: Accelerating Speculative Decoding for Visual Autoregressive Generation via Partial Verification SkippingHaotian Dong, Ye Li, Rongwei Lu, Chen Tang et al.CVPR 2026 · 4 citations
- Distilled Decoding 2: One-step Sampling of Image Auto-regressive Models with Conditional Score DistillationEnshu Liu, Qian Chen, Xuefei Ning, Shengen Yan et al.NeurIPS 2025 · 4 citations
Builds on18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
Related papers
- Grouped Speculative Decoding for Autoregressive Image GenerationJunhyuk So, Juncheol Shin, Hyunho Kook, Eunhyeok ParkICCV 2025 · 1 citation
- Spec-VLA: Speculative Decoding for Vision-Language-Action Models with Relaxed AcceptanceSongsheng Wang, Rucheng Yu, Zhihang Yuan, Chao Yu et al.EMNLP 2025
- Annealed Relaxation of Speculative Decoding for Faster Autoregressive Image GenerationXingyao Li, Fengzhuo Zhang, Cunxiao Du, Hui JiAAAI 2026
- Hawk: Leveraging Spatial Context for Faster Autoregressive Text-to-Image GenerationZhi-Kai Chen, Jun-Peng Jiang, Han-Jia Ye, De-Chuan ZhanNeurIPS 2025 · 2 citations
- See the Forest for the Trees: Loosely Speculative Decoding via Visual-Semantic Guidance for Efficient Inference of Video LLMsYicheng Ji, Jun Zhang, Jinpeng Chen, Cong Wang et al.ACL 2026 · 4 citations
