Efficient-SAM2: Accelerating SAM2 with Object-Aware Visual Encoding and Memory Retrieval
Jing Zhang, Zhikai Li, Xuewen Liu, Qingyi Gu
Abstract
Segment Anything Model 2 (SAM2) shows excellent performance in video object segmentation tasks; however, the heavy computational burden hinders its application in real-time video processing. Although there have been efforts to improve the efficiency of SAM2, most of them focus on retraining a lightweight backbone, with little exploration into post-training acceleration. In this paper, we observe that SAM2 exhibits sparse perception pattern as biological vision, which provides opportunities for eliminating redundant computation and acceleration: i) In mask decoder, the attention primarily focuses on the foreground objects, whereas the image encoder in the earlier stage exhibits a broad attention span, which results in unnecessary computation to background regions. ii) In memory bank, only a small subset of tokens in each frame contribute significantly to memory attention, and the salient regions exhibit temporal consistency, making full-token computation redundant. With these insights, we propose Efficient-SAM2, which promotes SAM2 to adaptively focus on object regions while eliminating task-irrelevant computations, thereby significantly improving inference efficiency. Specifically, for image encoder, we propose object-aware Sparse Window Routing (SWR), a window-level computation allocation mechanism that leverages the consistency and saliency cues from the previous-frame decoder to route background regions into a lightweight shortcut branch. Moreover, for memory attention, we propose object-aware Sparse Memory Retrieval (SMR), which allows only the salient memory tokens in each frame to participate in computation, with the saliency pattern reused from their first recollection. With negligible additional parameters and minimal training overhead, Efficient-SAM2 delivers 1.68× speedup on SAM2.1-L model with only 1.0% accuracy drop on SA-V test set, where SWR and SMR provide 1.83× and 1.78× speedups, respectively. Code is available at: https://github.com/jingjing0419/Efficient-SAM2 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 37e45522-ebf4-46da-b94a-aa31e5ac5914Builds on18
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- MViTv2: Improved Multiscale Vision Transformers for Classification and DetectionYanghao Li, Chao-Yuan Wu, Haoqi Fan, Karttikeya Mangalam et al.CVPR 2022 · 699 citations
- Hiera: A Hierarchical Vision Transformer without the Bells-and-WhistlesChaitanya Ryali, Yuan-Ting Hu, Daniel Bolya, Chen Wei et al.ICML 2023 · 388 citations
- MOSE: A New Dataset for Video Object Segmentation in Complex ScenesHenghui Ding, Chang Liu, Shuting He, Xudong Jiang et al.ICCV 2023 · 267 citations
- EfficientSAM: Leveraged Masked Image Pretraining for Efficient Segment AnythingYunyang Xiong, Bala Varadarajan, Lemeng Wu, Xiaoyu Xiang et al.CVPR 2024 · 185 citations
Related papers
- Efficient Track AnythingYunyang Xiong, Chong Zhou, Xiaoyu Xiang, Lemeng Wu et al.ICCV 2025 · 5 citations
- EdgeTAM: On-Device Track Anything ModelChong Zhou, Chenchen Zhu, Yunyang Xiong, Saksham Suri et al.CVPR 2025
- SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training CostHaiyang Mei, Pengyu Zhang, Mike Zheng ShouCVPR 2025
- Efficient Video Object Segmentation and Tracking with Recurrent Dynamic SubmodelWeidong Tang, Zhiyuan Liang, Xinyan Wan, Chen Zhu et al.CVPR 2026 · 2 citations
- Inter2Former: Dynamic Hybrid Attention for Efficient High-Precision Interactive SegmentationYou Huang, Lichao Chen, Jiayi Ji, Liujuan Cao et al.ICCV 2025 · 1 citation
