AdaSpark: Adaptive Sparsity for Efficient Long-Video Understanding
Handong Li, Zikang Liu, Longteng Guo, Tongtian Yue, Yepeng Tang, Xinxin Zhu, Chuanyang Zheng, Ziming Wang, Zhibin Wang, Jun Song, Cheng Yu, Bo Zheng, Jing Liu
Abstract
Processing long-form videos with Video Large Language Models (Video-LLMs) is computationally prohibitive. Current efficiency methods often compromise fine-grained perception through irreversible information disposal or inhibit long-range temporal modeling via rigid, predefined sparse patterns. This paper introduces AdaSpark, an adaptive sparsity framework designed to address these limitations. AdaSpark first partitions video inputs into 3D spatio-temporal cubes. It then employs two co-designed, context-aware components: (1) Adaptive Cube-Selective Attention (AdaS-Attn), which adaptively selects a subset of relevant video cubes to attend for each query token, and (2) Adaptive Token-Selective FFN (AdaS-FFN), which selectively processes only the most salient tokens within each cube. An entropy-based (Top-p) selection mechanism adaptively allocates computational resources based on input complexity. Experiments demonstrate that AdaSpark significantly reduces computational load by up to 57% FLOPs while maintaining comparable performance to dense models and preserving fine-grained, long-range dependencies, as validated on challenging hour-scale video benchmarks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fc3f502a-7e1e-4e00-904e-fd5a7a2e9ac9Builds on37
- Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse AttentionJingyang Yuan, Huazuo Gao, Damai Dai, Junyu Luo et al.ACL 2025 · 334 citations
- VAST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and DatasetSihan Chen, Handong Li, Qunbo Wang, Zijia Zhao et al.NeurIPS 2023 · 246 citations
- MoBA: Mixture of Block Attention for Long-Context LLMsEnzhe Lu, Zhejun Jiang, Jingyuan Liu, Yulun Du et al.NeurIPS 2025 · 219 citations
- VideoChat-Flash: Hierarchical Compression for Long-Context Video ModelingXinhao Li, Yi Wang, Jiashuo Yu, Xiangyu Zeng et al.ICLR 2026 · 172 citations
- Fit and Prune: Fast and Training-free Visual Token Pruning for Multi-modal Large Language ModelsWeihao Ye, Qiong Wu, Wenhao Lin, Yiyi ZhouAAAI 2025 · 99 citations
Related papers
- VecAttention: Vector-wise Sparse Attention for Accelerating Long Context InferenceAnmin Liu, Ruixuan Yang, Huiqiang Jiang, Bin Lin et al.CVPR 2026 · 4 citations
- Streaming Long Video Understanding with Large Language ModelsRui Qian, Xiaoyi Dong, Pan Zhang, Yuhang Zang et al.NeurIPS 2024 · 216 citations
- LongVU: Spatiotemporal Adaptive Compression for Long Video-Language UnderstandingXiaoqian Shen, Yunyang Xiong, Changsheng Zhao, Lemeng Wu et al.ICML 2025
- Keyframe-Oriented Vision Token Pruning: Enhancing Efficiency of Large Vision Language Models on Long-form Video ProcessingYudong Liu, Jingwei Sun, Yueqian Lin, Jianyi Zhang et al.ICCV 2025 · 22 citations
- FlexSelect: Flexible Token Selection for Efficient Long Video UnderstandingYunzhu Zhang, Yu Lu, Tianyi Wang, Fengyun Rao et al.NeurIPS 2025 · 22 citations
