Efficient Video Object Segmentation and Tracking with Recurrent Dynamic Submodel
Weidong Tang, Zhiyuan Liang, Xinyan Wan, Chen Zhu, Zhaopan Xu, Pengfei Zhou, Yan Song, Yang You, Wangbo Zhao
Abstract
Large vision foundation models, such as SAM2, have achieved remarkable performance in video object segmentation and tracking (VOST). However, their effectiveness is hindered by significant computational overhead. While model pruning is a widely used strategy to address this issue, traditional static and input-agnostic pruning approaches fall short in managing the diverse and complex nature of video data effectively. A promising alternative is dynamic networks, yet they often struggle to translate theoretical computational reductions into actual acceleration. Furthermore, both static and dynamic approaches typically focus on visual features of individual frames while neglecting the temporal correlations between them, limiting their performance in handling complex video streams. To address these challenges, we propose Recurrent Dynamic Submodel (RDS), a dynamic architecture that adaptively selects submodel blocks for each frame. Specifically, it has a lightweight Prediction-Aware Router (PAR), which leverages both the segmentation mask from the previous frame and the visual features of the current frame to make routing decisions, enabling the submodel to explicitly capture the temporal nature of video data. Additionally, to reduce the cost of adapting the dynamic submodel, we introduce an Importance-aware LoRA (I-LoRA), tuning parameters only in the most critical blocks. Extensive experiments show that our approach achieves an optimal trade-off between performance and speed, requiring minimal trainable parameters and training data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on18
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- DynamicViT: Efficient Vision Transformers with Dynamic Token SparsificationYongming Rao, Wenliang Zhao, Benlin Liu, Jiwen Lu et al.NeurIPS 2021 · 1,343 citations
- Self Forcing: Bridging the Train-Test Gap in Autoregressive Video DiffusionXun Huang, Zhengqi Li, Guande He, Mingyuan Zhou et al.NeurIPS 2025 · 628 citations
- A-ViT: Adaptive Tokens for Efficient Vision TransformerHongxu Yin, Arash Vahdat, José M. Álvarez, Arun Mallya et al.CVPR 2022 · 288 citations
- Streaming Long Video Understanding with Large Language ModelsRui Qian, Xiaoyi Dong, Pan Zhang, Yuhang Zang et al.NeurIPS 2024 · 216 citations
Related papers
- Efficient-SAM2: Accelerating SAM2 with Object-Aware Visual Encoding and Memory RetrievalJing Zhang, Zhikai Li, Xuewen Liu, Qingyi GuICLR 2026 · 5 citations
- Parameter-Free Fine-tuning via Redundancy Elimination for Vision Foundation ModelsJiahuan Long, Tingsong Jiang, Wen Yao, Yizhe Xiong et al.AAAI 2026
- LoRATv2: Enabling Low-Cost Temporal Modeling in One-Stream TrackersLiting Lin, Heng Fan, Zhipeng Zhang, Yuqing Huang et al.NeurIPS 2025 · 11 citations
- SAM2Text: Towards Prompt-Free and Multi-Resolution Video Scene Text SegmentationJing-Yao Zhang, Heng Zhang, Mingsen Zhang, Binbin Yang et al.CVPR 2026
- Efficient Track AnythingYunyang Xiong, Chong Zhou, Xiaoyu Xiang, Lemeng Wu et al.ICCV 2025 · 5 citations
