SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation
Claudia Cuttano, Gabriele Trivigno, Gabriele Rosi, Carlo Masone, Giuseppe Averta
Abstract
Require entire video Context propagation TARGET: absent MATCH: red shirt LEARNABLE CORRECTION: detected object switch Query: 'The cyclist in red that overtakes others on the left' MOTION REASONING: action that unfolds over multiple frames TRACKING: occlusion handling Offline + Efficient processing -No global context Clip-based Figure 1. SAMWISE. Our approach infuses knowledge about natural language in the Segment-Anything 2 model, adding explicit temporal cues in the feature extraction for the task of streaming-based Referring Video Segmentation (RVOS). We use a learnable mechanism to mitigate the so-called tracking bias, i.e. SAM2 tendency to overlook a correct object once it becomes identifiable, due to its ongoing tracking of a different object. Our design enables effective streaming processing for RVOS, exploiting the memory from previous frames to propagate past context. The figure shows an example where the target object is not present in the first frame, leading SAM2 to start tracking the wrong one. Afterwards, when the correct object appears, our learnable correction mechanisms guides SAM2 to switch its tracking focus. By adding in its features the notion of temporal evolution, the model is able to recognize that the new object is more aligned with the provided textual query. Finally, we exploit SAM2 tracking skills and robustness to occlusions to keep following the object.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 30bf301f-46ce-446b-b110-9fb8e5dfd0b1Cited by top-tier papers1
Ask how each one uses itBuilds on26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
Related papers
- MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object SegmentationFu Rong, Meng Lan, Qian Zhang, Lefei ZhangICCV 2025 · 4 citations
- Correspondence as Video: Test-Time Adaption on SAM2 for Reference Segmentation in the WildHaoran Wang, Zekun Li, Jian Zhang, Lei Qi et al.ICCV 2025
- Towards Streaming Referring Video Segmentation via Large Language ModelWenkang Zhang, Kaicheng Yang, Xiang An, Qiang Li et al.CVPR 2026
- SAM2LONG: Enhancing SAM 2 for Long Video Segmentation with a Training-Free Memory TreeShuangrui Ding, Rui Qian, Xiaoyi Dong, Pan Zhang et al.ICCV 2025 · 15 citations
- LoSh: Long-Short Text Joint Prediction Network for Referring Video Object SegmentationLinfeng Yuan, Miaojing Shi, Zijie Yue, Qijun ChenCVPR 2024 · 12 citations
