Weakly Supervised Video Salient Object Detection via Point Supervision
Shuyong Gao, Haozhe Xing, Wei Zhang, Yan Wang, Qianyu Guo, Wenqiang Zhang
Abstract
Fully supervised video salient object detection models have achieved excellent performance, yet obtaining pixel-by-pixel annotated datasets is laborious. Several works attempt to use scribble annotations to mitigate this problem, but point supervision as a more labor-saving annotation method (even the most labor-saving method among manual annotation methods for dense prediction), has not been explored. In this paper, we propose a strong baseline model based on point supervision. To infer saliency maps with temporal information, we mine inter-frame complementary information from short-term and long-term perspectives, respectively. Specifically, we propose a hybrid token attention module, which mixes optical flow and image information from orthogonal directions, adaptively highlighting critical optical flow information (channel dimension) and critical token information (spatial dimension). To exploit long-term cues, we develop the Long-term Cross-Frame Attention module (LCFA), which assists the current frame in inferring salient objects based on multi-frame tokens. Furthermore, we label two point-supervised datasets, P-DAVIS and P-DAVSOD, by relabeling the DAVIS and the DAVSOD dataset. Experiments on the six benchmark datasets illustrate our method outperforms the previous state-of-the-art weakly supervised methods and even is comparable with some fully supervised approaches. Our source code and datasets are available at: https://github.com/shuyonggao/PVSOD.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 38b2f1e9-3e63-469d-9426-ef2bcdfeb29aCited by top-tier papers3
- ZOOM: Learning Video Mirror Detection with Extremely-Weak SupervisionKe Xu, Tsun Wai Siu, Rynson W. H. LauAAAI 2024 · 10 citations
- Scoring, Remember, and Reference: Catching Camouflaged Objects in VideosYu'ang Feng, Shuyong Gao, Fuzhen Yan, Yicheng Song et al.ICCV 2025 · 2 citations
- Bridging RGB and Hematoxylin Components: An Interleaved Guidance and Fusion Framework for Point Supervised Nuclei SegmentationZihan Huan, Xipeng Pan, Hualong Zhang, Siyang Feng et al.CVPR 2026
Builds on23
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan et al.ICCV 2021 · 4,909 citations
- Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNetLi Yuan, Yunpeng Chen, Tao Wang, Weihao Yu et al.ICCV 2021 · 2,462 citations
Related papers
- Weakly Supervised Video Salient Object DetectionWangbo Zhao, Jing Zhang, Long Li, Nick Barnes et al.CVPR 2021
- Weakly-Supervised Salient Object Detection Using Point SupervisonShuyong Gao, Wei Zhang, Yan Wang, Qianyu Guo et al.AAAI 2022 · 77 citations
- Semi-Supervised Video Salient Object Detection Using Pseudo-LabelsPengxiang Yan, Guanbin Li, Yuan Xie, Zhen Li et al.ICCV 2019 · 134 citations
- Motion Guided Attention for Video Salient Object DetectionHaofeng Li, Guanqi Chen, Guanbin Li, Yizhou YuICCV 2019 · 200 citations
- Weakly-Supervised Salient Object Detection via Scribble AnnotationsJing Zhang, Xin Yu, Aixuan Li, Peipei Song et al.CVPR 2020
