SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost
Haiyang Mei, Pengyu Zhang, Mike Zheng Shou
摘要
Foundation models like the Segment Anything Model (SAM) have significantly advanced promptable image segmentation in computer vision. However, extending these capabilities to videos presents substantial challenges, particularly in ensuring precise and temporally consistent mask propagation in dynamic scenes. SAM 2 attempts to address this by training a model on massive image and video data from scratch to learn complex spatiotemporal associations, resulting in huge training costs that hinder research and practical deployment. In this paper, we introduce SAM-I2V, an effective image-to-video upgradation method for cultivating a promptable video segmentation (PVS) model. Our approach strategically upgrades the pretrained SAM to support PVS, significantly reducing training complexity and resource requirements. To achieve this, we introduce three key innovations: (i) an image-to-video feature extraction upgrader built upon SAM's static image encoder to enable spatiotemporal video perception, (ii) a memory filtering strategy that selects the most relevant past frames for more effective utilization of historical information, and (iii) a memory-as-prompt mechanism leveraging object memory to ensure temporally consistent mask propagation in dynamic scenes. Comprehensive experiments demonstrate that our method achieves over 90% of SAM 2's performance while using only 0.2% of its training cost. Our work presents a resource-efficient pathway to PVS, lowering barriers for further research in PVS model design and enabling broader applications and advancements in the field.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- RobotSeg: A Model and Dataset for Segmenting Robots in Image and VideoHaiyang Mei, Qiming Huang, Hai Ci, Mike Zheng ShouCVPR 2026 · 被引用 3 次
- VIRST: Video-Instructed Reasoning Assistant for SpatioTemporal SegmentationJihwan Hong, Jaeyoung DoCVPR 2026 · 被引用 2 次
- Robust Promptable Video Object SegmentationSohyun Lee, Yeho Gwon, Lukas Hoyer, Konrad Schindler 等CVPR 2026
- PolarDepth: Monocular Transparent Object Depth from Polar-Physics PriorsWen Dong, Haiyang Mei, Yinglian Ji, Zijun Zhang 等ICML 2026
- Temporal-Emerged Prompting for Segment Anything in Multiframe Infrared Small Target DetectionYinghui Xing, Donghao Chu, Shizhou Zhang, di xuICML 2026
它引用的顶会 Paper20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- AdaptFormer: Adapting Vision Transformers for Scalable Visual RecognitionShoufa Chen, Chongjian Ge, Zhan Tong, Jiangliu Wang 等NeurIPS 2022 · 被引用 1,291 次
- Segment Anything in High QualityLei Ke, Mingqiao Ye, Martin Danelljan, Yifan Liu 等NeurIPS 2023 · 被引用 709 次
相关 Paper
- OFL-SAM2: Prompt SAM2 with Online Few-shot Learner for Efficient Medical Image SegmentationMeng Lan, Lefei Zhang, Xiaomeng LiAAAI 2026
- MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object SegmentationFu Rong, Meng Lan, Qian Zhang, Lefei ZhangICCV 2025 · 被引用 4 次
- Endow SAM with Keen Eyes: Temporal-Spatial Prompt Learning for Video Camouflaged Object DetectionWenjun Hui, Zhenfeng Zhu, Shuai Zheng, Yao ZhaoCVPR 2024
- Efficient Track AnythingYunyang Xiong, Chong Zhou, Xiaoyu Xiang, Lemeng Wu 等ICCV 2025 · 被引用 5 次
- Towards Fine-Grained Interactive Segmentation in Images and VideosYuan Yao, Qiushi Yang, Miaomiao Cui, Liefeng BoICCV 2025 · 被引用 2 次
