VIA-SD: Verification via Intra-Model Routing for Speculative Decoding
Yuchen Xian, Yang He, Yunqiu Xu, Yi Yang
摘要
Speculative decoding (SD) addresses the high inference costs of LLMs by having lightweight drafters generate candidates for large verifiers to validate in parallel. Existing draft-verify methods use binary decisions: accept or fully recompute. Yet we find that many rejected tokens can be verified correctly by a slim submodel derived from the full verifier via intra-model routing, instead of the full verifier. This motivates our slim-verifier to handle tokens requiring moderate verification resources, reducing expensive large-model calls. We propose V erification via I ntr a -Model Routing for S peculative D ecoding (VIA-SD), a multi-tier framework using a routed slim-verifier. Draft tokens are processed hierarchically: direct acceptance for high-confidence cases, slim-verifier regeneration for medium-confidence cases, and full-model verification for uncertain cases. Across four representative tasks and multiple model families, VIA-SD reduces rejection rates by 0.10–0.22 and delivers 10–20% speedups over strong SD baselines, while achieving 2.5–3× acceleration over non-drafting decoding. Moreover, VIA-SD is compatible with existing SD frameworks without modifying their training procedures. Our results suggest multi-tier SD as a general paradigm for scalable and efficient LLM inference. Project page: https://zju-xyc.github.io/VIA-SD-Project-Page/
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper15
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 被引用 1,472 次
- Efficient large-scale language model training on GPU clusters using megatron-LMDeepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley 等SC 2021 · 被引用 576 次
- Coresets for Data-efficient Training of Machine Learning ModelsBaharan Mirzasoleiman, Jeff A. Bilmes, Jure LeskovecICML 2020 · 被引用 494 次
- EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time TestYuhui Li, Fangyun Wei, Chao Zhang, Hongyang ZhangNeurIPS 2025 · 被引用 347 次
- Speculative Decoding with Big Little DecoderSehoon Kim, Karttikeya Mangalam, Suhong Moon, Jitendra Malik 等NeurIPS 2023 · 被引用 212 次
相关 Paper
- SWIFT: On-the-Fly Self-Speculative Decoding for LLM Inference AccelerationHeming Xia, Yongqi Li, Jun Zhang, Cunxiao Du 等ICLR 2025
- MineDraft: A Framework for Batch Parallel Speculative DecodingZhenwei Tang, Arun Verma, Zijian Zhou, Zhaoxuan Wu 等ICML 2026
- Speculative Decoding with CTC-based Draft Model for LLM Inference AccelerationZhuofan Wen, Shangtong Gui, Yang FengNeurIPS 2024 · 被引用 19 次
- Draft& Verify: Lossless Large Language Model Acceleration via Self-Speculative DecodingJun Zhang, Jue Wang, Huan Li, Lidan Shou 等ACL 2024
- SelfJudge: Faster Speculative Decoding via Self-Supervised Judge VerificationKanghoon Yoon, Minsub Kim, Sungjae Lee, Joonhyung Lee 等ICML 2026 · 被引用 4 次
