Sparsity Forcing: Reinforcing Token Sparsity of MLLMs
Feng Chen, Yefei He, Lequan Lin, Jing Liu, Chenhui Gou, Bohan Zhuang, Qi Wu
摘要
Sparse attention mechanisms aim to reduce computational overhead with minimal accuracy loss by selectively processing salient tokens. Despite their effectiveness, most methods merely exploit a model’s inherent sparsity and thus plateau at moderate budgets (about 50% token reduction), with little headroom to push budget lower without hurting accuracy. Other approaches attempt to enforce sparsity through trainable sparse attention or sharpness-inducing regularizers, but these either fix rigid patterns that ignore input and layer dynamics, or optimize proxy objectives without direct control over token budgets. In this paper, we explicitly reinforce token sparsity in well-posed multimodal large language models (MLLMs) through a simple RL-based post-training framework named . Our method explores the efficiency-accuracy trade-off by running multiple rollouts with different token budgets, where both efficiency (token reduction ratio) and performance (answer correctness) are formulated as joint rewards. By contrasting rollouts within each group, the more efficient and correct answer is rewarded while less efficient or incorrect ones are penalized, thereby turning token saving into an end-to-end, inference-consistent optimization objective. Across thirteen image and video benchmarks, Sparsity Forcing raises token reduction ratio on Qwen2-VL/Qwen2.5-VL from 20% to 75% with minimal accuracy decline, significantly reducing long-context inference memory by up to 3 while speeding up decoding by up to 3.3.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper15
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan 等NeurIPS 2025 · 被引用 2,828 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- Video-R1: Reinforcing Video Reasoning in MLLMsKaituo Feng, Kaixiong Gong, Bohao Li, Zonghao Guo 等NeurIPS 2025 · 被引用 528 次
相关 Paper
- SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMsJiahui Wang, Zuyan Liu, Yongming Rao, Jiwen LuICCV 2025 · 被引用 1 次
- OmniSparse: Training-Aware Fine-Grained Sparse Attention for Long-Video MLLMsFeng Chen, Yefei He, Shaoxuan He, Yuanyu He 等AAAI 2026
- SparseVILA: Decoupling Visual Sparsity for Efficient VLM InferenceSamir Khaki, Junxian Guo, Jiaming Tang, Shang Yang 等ICCV 2025 · 被引用 3 次
- ZipVL: Accelerating Vision-Language Models Through Dynamic Token SparsityYefei He, Feng Chen, Jing Liu, Wenqi Shao 等ICCV 2025 · 被引用 1 次
- MARC: Memory-Augmented RL Token Compression for Efficient Video UnderstandingPeiran Wu, Zhuorui Yu, Yunze Liu, Chi-Hao Wu 等ICLR 2026 · 被引用 7 次
