Lune

CVPR2026Top-tier venue

ParallelVLM: Lossless Video-LLM Acceleration with Visual Alignment Aware Parallel Speculative Decoding

Quan Kong, Yuhao Shen, Yicheng Ji, Huan Li, Cong Wang

2026Year
7Citations
2Top-tier citations

Abstract

Although current Video-LLMs achieve impressive performance in video understanding tasks, their autoregressive decoding efficiency remains constrained by the massive number of video tokens. Visual token pruning can partially ease this bottleneck, yet existing approaches still suffer from information loss and yield only modest acceleration in decoding. In this paper, we propose ParallelVLM, a training-free draft-then-verify speculative decoding framework that overcomes both mutual waiting and limited speedup-ratio problems between draft and target models in long-video settings. ParallelVLM features two parallelized stages that maximize hardware utilization and incorporates an Unbiased Verifier-Guided Pruning strategy to better align the draft and target models by eliminating the positional bias in attention‑guided pruning. Extensive experiments demonstrate that ParallelVLM effectively expands the draft window by 1.6∼1.8×1.6\sim1.8\times with high accepted lengths, and accelerates various video understanding benchmarks by 3.36×\times on LLaVA-Onevision-72B and 2.42×\times on Qwen2.5-VL-32B compared with vanilla autoregressive decoding.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 798107a0-757e-4988-aab5-7a6b2ff7ce99

Cited by top-tier papers2

Ask how each one uses it

Builds on33

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines