Dual Prototype Attention for Unsupervised Video Object Segmentation
Suhwan Cho, Minhyeok Lee, Seunghoon Lee, Dogyoon Lee, Heeseung Choi, Ig-Jae Kim, Sangyoun Lee
Abstract
Unsupervised video object segmentation (VOS) aims to detect and segment the most salient object in videos. The primary techniques used in unsupervised VOS are 1) the collaboration of appearance and motion information; and 2) temporal fusion between different frames. This paper proposes two novel prototype-based attention mechanisms, inter-modality attention (IMA) and inter-frame attention (IFA), to incorporate these techniques via dense propagation across different modalities and frames. IMA densely integrates context information from different modalities based on a mutual refinement. IFA injects global context of a video to the query frame, enabling a full utilization of useful properties from multiple frames. Experimental results on public benchmark datasets demonstrate that our proposed approach outperforms all existing methods by a substantial margin. The proposed two components are also thoroughly validated via ablative study. Code and models are available at https://github.com/Hydragon516/DPA .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Matting Anything 2: Towards Video Matting for AnythingChenyi Zhang, Yiheng Lin, Yunchao Wei, Hongsong Wang et al.ICLR 2026
- M4-SAM: Multi-Modal Mixture-of-Experts with Memory-Augmented SAM for RGB-D Video Salient Object DetectionJiyuan Liu, Jia Lin, Xiaofei Zhou, Runmin Cong et al.CVPR 2026
- Segment Any Motion in VideosNan Huang, Wenzhao Zheng, Chenfeng Xu, Kurt Keutzer et al.CVPR 2025
- SAM-DAQ: Segment Anything Model with Depth-guided Adaptive Queries for RGB-D Video Salient Object DetectionJia Lin, Xiaofei Zhou, Jiyuan Liu, Runmin Cong et al.AAAI 2026
- OSMamba: Omnidirectional Spectral Mamba with Dual-Domain Prior Generator for Exposure CorrectionGehui Li, Bin Chen, Chen Zhao, Lei Zhang et al.CVPR 2025
Builds on10
- Zero-Shot Video Object Segmentation via Attentive Graph Neural NetworksWenguan Wang, Xiankai Lu, Jianbing Shen, David J. Crandall et al.ICCV 2019 · 294 citations
- Motion-Attentive Transition for Zero-Shot Video Object SegmentationTianfei Zhou, Shunzhou Wang, Yi Zhou, Yazhou Yao et al.AAAI 2020 · 210 citations
- Full-Duplex Strategy for Video Object SegmentationGe-Peng Ji, Keren Fu, Zhe Wu, Deng-Ping Fan et al.ICCV 2021 · 173 citations
- Anchor Diffusion for Unsupervised Video Object SegmentationZhao Yang, Qiang Wang, Luca Bertinetto, Song Bai et al.ICCV 2019 · 127 citations
- Learning Motion-Appearance Co-Attention for Zero-Shot Video Object SegmentationShu Yang, Lu Zhang, Jinqing Qi, Huchuan Lu et al.ICCV 2021 · 76 citations
Related papers
- SimulFlow: Simultaneously Extracting Feature and Identifying Target for Unsupervised Video Object SegmentationLingyi Hong, Wei Zhang, Shuyong Gao, Hong Lu et al.ACM MM 2023 · 14 citations
- Guided Slot Attention for Unsupervised Video Object SegmentationMinhyeok Lee, Suhwan Cho, Dogyoon Lee, Chaewon Park et al.CVPR 2024
- Semantics Meets Temporal Correspondence: Self-supervised Object-centric Learning in VideosRui Qian, Shuangrui Ding, Xian Liu, Dahua LinICCV 2023 · 23 citations
- Multi-Source Fusion and Automatic Predictor Selection for Zero-Shot Video Object SegmentationXiaoqi Zhao, Youwei Pang, Jiaxing Yang, Lihe Zhang et al.ACM MM 2021 · 35 citations
- Shallow Features Matter: Hierarchical Memory with Heterogeneous Interaction for Unsupervised Video Object SegmentationXiangyu Zheng, Songcheng He, Wanyun Li, Xiaoqiang Li et al.ACM MM 2025
