Multi-Scale Local Speculative Decoding for Image Generation
Elia Peruzzo, Guillaume Sautiere, Amir Habibian
摘要
Autoregressive (AR) models have achieved remarkable success in image synthesis, yet their sequential nature imposes significant latency constraints. Speculative Decoding offers a promising avenue for acceleration, but existing approaches are limited by token-level ambiguity and lack of spatial awareness. In this work, we introduce Multi-Scale Local Speculative Decoding (MULO-SD), a novel framework that combines multi-resolution drafting with spatially informed verification to accelerate AR image generation. Our method leverages a low-resolution drafter paired with an up-sampling step to propose candidate image tokens, which are then verified in parallel by a high-resolution target model. Crucially, we incorporate a local rejection and resampling mechanism, enabling efficient correction of draft errors by focusing on spatial neighborhoods rather than raster-scan resampling after the first rejection. When integrated with parallel decoding resampling, MULO-SD achieves substantial speedups -up to 5× -outperforming both speculative decoding and parallel decoding baselines in terms of acceleration, while maintaining comparable semantic alignment and perceptual quality. These results are validated using GenEval, DPG-Bench, and FID/HPSv2 on the MS-COCO 5k validation split. Extensive ablations highlight the impact of up-sampling design, probability pooling, and local rejection and resampling with neighborhood expansion. Our approach sets a new state-of-the-art in speculative decoding for image synthesis, bridging the gap between efficiency and fidelity. Project page is available at https://qualcomm-ai- research.github.io/mulo-sd-webpage/.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper15
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann 等ICLR 2024 · 被引用 4,569 次
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 被引用 1,472 次
- Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale PredictionKeyu Tian, Yi Jiang, Zehuan Yuan, Bingyue Peng 等NeurIPS 2024 · 被引用 1,199 次
- Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned RepresentationsJiaming Han, Hao Chen, Yang Zhao, Hanyu Wang 等NeurIPS 2025 · 被引用 50 次
- Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache CompressionKunjun Li, Zigeng Chen, Cheng-Yen Yang, Jenq-Neng HwangNeurIPS 2025 · 被引用 23 次
相关 Paper
- VVS: Accelerating Speculative Decoding for Visual Autoregressive Generation via Partial Verification SkippingHaotian Dong, Ye Li, Rongwei Lu, Chen Tang 等CVPR 2026 · 被引用 4 次
- Towards Better & Faster Autoregressive Image Generation: From the Perspective of EntropyXiaoxiao Ma, Feng Zhao, Pengyang Ling, Haibo Qiu 等NeurIPS 2025 · 被引用 12 次
- Locality-aware Parallel Decoding for Efficient Autoregressive Image GenerationZhuoyang Zhang, Luke J. Huang, Chengyue Wu, Shang Yang 等ICLR 2026 · 被引用 8 次
- Annealed Relaxation of Speculative Decoding for Faster Autoregressive Image GenerationXingyao Li, Fengzhuo Zhang, Cunxiao Du, Hui JiAAAI 2026
- Speculative Speculative DecodingTanishq Kumar, Tri Dao, Avner MayICLR 2026 · 被引用 15 次
