Lune

ICLR2026Top-tier venue

SURGE: Surprise-Guided Token Reduction for Efficient Video Understanding with VLMs

Chong Tang, Sannara Ek, Dirk Koch, Robert Mullins, Alex S. Weddell, Jagmohan Chauhan

2026Year

Abstract

Videos contain rich information but also high redundancy, as consecutive frames often share similar backgrounds and predictable motions. Current video-language models (VLMs) are unable to exploit this redundancy and therefore perform a significant amount of superfluous computation, processing thousands of patch tokens even when little new information is present. What is missing is an onthe-fly, model-agnostic signal of temporal predictability to decide whether tokens carry unpredictable information that merits computation. We propose SURGE, a training-free and backbone-agnostic method that measures surprise in token space. Surprise scores are defined by the prediction error of each token from its recent history; high-surprise tokens are retained, while predictable ones are pruned. Aggregating scores over time produces a surprise curve that highlights key events, which can be further refined with CLIP-based query relevance to form a compact spatio-temporal mask. Experiments on multiple video understanding benchmarks show that SURGE reduces tokens by up to 7× and prefill cost by 86-98%, while maintaining accuracy within ±1 point of full-token baselines. By aligning computation with novelty, SURGE enables video VLMs to handle long contexts efficiently and without retraining. https://github.com/BarryTang22/ SURGE.git

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 005f68fb-9375-49a6-953d-8dae1b903313

Builds on30

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines