LANTERN: Accelerating Visual Autoregressive Models with Relaxed Speculative Decoding
Doohyuk Jang, Sihwan Park, June Yong Yang, Yeonsung Jung, Jihun Yun, Souvik Kundu, Sungyub Kim, Eunho Yang
摘要
Auto-Regressive (AR) models have recently gained prominence in image generation, often matching or even surpassing the performance of diffusion models. However, one major limitation of AR models is their sequential nature, which processes tokens one at a time, slowing down generation compared to models like GANs or diffusion-based methods that operate more efficiently. While speculative decoding has proven effective for accelerating LLMs by generating multiple tokens in a single forward, its application in visual AR models remains largely unexplored. In this work, we identify a challenge in this setting, which we term token selection ambiguity, wherein visual AR models frequently assign uniformly low probabilities to tokens, hampering the performance of speculative decoding. To overcome this challenge, we propose a relaxed acceptance condition referred to as LANTERN that leverages the interchangeability of tokens in latent space. This relaxation restores the effectiveness of speculative decoding in visual AR models by enabling more flexible use of candidate tokens that would otherwise be prematurely rejected. Furthermore, by incorporating a total variation distance bound, we ensure that these speed gains are achieved without significantly compromising image quality or semantic coherence. Experimental results demonstrate the efficacy of our method in providing a substantial speed-up over speculative decoding. In specific, compared to a naïve application of the state-of-the-art speculative decoding, LANTERN increases speed-ups by 1.75× and 1.82×, as compared to greedy decoding and random sampling, respectively, when applied to LlamaGen, a contemporary visual AR model. The code is publicly available at https://github.com/jadohu/LANTERN .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper22
- Towards Better & Faster Autoregressive Image Generation: From the Perspective of EntropyXiaoxiao Ma, Feng Zhao, Pengyang Ling, Haibo Qiu 等NeurIPS 2025 · 被引用 12 次
- Speculative Jacobi-Denoising Decoding for Accelerating Autoregressive Text-to-image GenerationYao Teng, Fuyun Wang, Xian Liu, Zhekai Chen 等NeurIPS 2025 · 被引用 8 次
- FreqExit: Enabling Early-Exit Inference for Visual Autoregressive Models via Frequency-Aware GuidanceYing Li, Chengfei Lyu, Huan WangNeurIPS 2025 · 被引用 6 次
- VVS: Accelerating Speculative Decoding for Visual Autoregressive Generation via Partial Verification SkippingHaotian Dong, Ye Li, Rongwei Lu, Chen Tang 等CVPR 2026 · 被引用 4 次
- Distilled Decoding 2: One-step Sampling of Image Auto-regressive Models with Conditional Score DistillationEnshu Liu, Qian Chen, Xuefei Ning, Shengen Yan 等NeurIPS 2025 · 被引用 4 次
它引用的顶会 Paper18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
相关 Paper
- Grouped Speculative Decoding for Autoregressive Image GenerationJunhyuk So, Juncheol Shin, Hyunho Kook, Eunhyeok ParkICCV 2025 · 被引用 1 次
- Spec-VLA: Speculative Decoding for Vision-Language-Action Models with Relaxed AcceptanceSongsheng Wang, Rucheng Yu, Zhihang Yuan, Chao Yu 等EMNLP 2025
- Annealed Relaxation of Speculative Decoding for Faster Autoregressive Image GenerationXingyao Li, Fengzhuo Zhang, Cunxiao Du, Hui JiAAAI 2026
- Hawk: Leveraging Spatial Context for Faster Autoregressive Text-to-Image GenerationZhi-Kai Chen, Jun-Peng Jiang, Han-Jia Ye, De-Chuan ZhanNeurIPS 2025 · 被引用 2 次
- See the Forest for the Trees: Loosely Speculative Decoding via Visual-Semantic Guidance for Efficient Inference of Video LLMsYicheng Ji, Jun Zhang, Jinpeng Chen, Cong Wang 等ACL 2026 · 被引用 4 次
