Lune

ACM MM2025Top-tier venue

VSumMamba: Mamba Empowered Efficient Video Summarization with Multi-Scale Spatial-Temporal Modeling

Yamiao Ding, Tianrui Liu, Zhizhou Lu, Jun-Jie Huang, Wentao Zhao, Xinwang Liu, Meng Wang

2025Year
1Citations

Abstract

The exponential growth of video content necessitates efficient summarization techniques that balance local redundancy reduction and global dependency modeling. In this work, we introduce VSumMamba, an innovative video summarization approach that leverages Selective State Space Models to address the quadratic complexity limitations of Transformer based approaches meanwhile surpassing CNNs' restricted long-range modeling capabilities. The proposed framework comprises three core components: 1) a Multi-Scale Aggregator, 2) a Cascaded Temporal Modeling Module with bi-directional Mamba blocks for temporal representation enhancement, and 3) a Parallel Spatial Modeling Module employing spatial Mamba blocks, operating in concert to effectively refine spatiotemporal video representations. Through three specialized multi-scale spatial-temporal modeling schemes, VSumMamba demonstrate the ability to balance computational efficiency and summarization performance. Comprehensive evaluations on benchmarks datasets demonstrate VSumMamba's superior performance, achieving 67.5% and 56.0% F1-scores on TVSum and SumMe respectively, while maintaining lower computational cost compared to existing state-of-the-art methods.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

lune papers get 99542ca0-67bd-4178-82bb-dfdcfd0df44b

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines