AASD: Accelerate Inference by Aligning Speculative Decoding in Multimodal Large Language Models
Chaoqun Yang, Ran Chen, Muyang Zhang, Weiguang Pang, Yuzhi Chen, Rongtao Xu, Kexue Fu, Changwei Wang, Longxiang Gao
Abstract
Multimodal Large Language Models (MLLMs) have achieved notable success in visual instruction tuning, yet their inference is time-consuming due to the auto-regressive decoding of Large Language Model (LLM) backbone. Traditional methods for accelerating inference, including model compression and migration from language model acceleration, often compromise output quality or face challenges in effectively integrating multimodal features. To address these issues, we propose AASD, a novel framework for Accelerating inference with refined KV Cache and Aligning speculative decoding in MLLMs. Our approach leverages the target model’s cached KeyValue (KV) pairs to extract vital information for generating draft tokens, enabling efficient speculative decoding. To reduce the computational burden associated with long multimodal token sequences, we introduce a KV Projector to compress the KV Cache while maintaining representational fidelity. Additionally, we design a Target-Draft Attention mechanism that optimizes the alignment between the draft model and the target model, achieving the benefits of real inference scenarios with minimal computational overhead. Extensive experiments on mainstream MLLMs demonstrate that our method achieves up to a inference speedup without sacrificing accuracy. This study not only provides an effective and lightweight solution for accelerating MLLM inference but also introduces a novel alignment strategy for speculative decoding in multimodal contexts, laying a strong foundation for future research in efficient MLLMs. Code is availiable at https://github.com/transcend-0/ASD
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 4e62ab5e-0cbe-4f51-a8d9-ecb95902f95dRelated papers
- ViSpec: Accelerating Vision-Language Models with Vision-Aware Speculative DecodingJialiang Kang, Han Shu, Wenshuo Li, Yingjie Zhai et al.NeurIPS 2025 · 24 citations
- DREAM: Drafting with Refined Target Features and Entropy-Adaptive Cross-Attention Fusion for Multimodal Speculative DecodingYunhai Hu, Tianhua Xia, Zining Liu, Rahul Raman et al.NeurIPS 2025 · 15 citations
- VL-Cache: Sparsity and Modality-Aware KV Cache Compression for Vision-Language Model Inference AccelerationDezhan Tu, Danylo Vashchilenko, Yuzhe Lu, Panpan XuICLR 2025
- SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMsJiahui Wang, Zuyan Liu, Yongming Rao, Jiwen LuICCV 2025 · 1 citation
- Revisiting Multimodal KV Cache Compression: A Frequency-Domain-Guided Outlier-KV-Aware ApproachYaoxin Yang, Peng Ye, Xudong Tan, Chongjun Tu et al.CVPR 2026 · 5 citations
