Weakly Supervised Video Salient Object Detection via Point Supervision
Shuyong Gao, Haozhe Xing, Wei Zhang, Yan Wang, Qianyu Guo, Wenqiang Zhang
摘要
Fully supervised video salient object detection models have achieved excellent performance, yet obtaining pixel-by-pixel annotated datasets is laborious. Several works attempt to use scribble annotations to mitigate this problem, but point supervision as a more labor-saving annotation method (even the most labor-saving method among manual annotation methods for dense prediction), has not been explored. In this paper, we propose a strong baseline model based on point supervision. To infer saliency maps with temporal information, we mine inter-frame complementary information from short-term and long-term perspectives, respectively. Specifically, we propose a hybrid token attention module, which mixes optical flow and image information from orthogonal directions, adaptively highlighting critical optical flow information (channel dimension) and critical token information (spatial dimension). To exploit long-term cues, we develop the Long-term Cross-Frame Attention module (LCFA), which assists the current frame in inferring salient objects based on multi-frame tokens. Furthermore, we label two point-supervised datasets, P-DAVIS and P-DAVSOD, by relabeling the DAVIS and the DAVSOD dataset. Experiments on the six benchmark datasets illustrate our method outperforms the previous state-of-the-art weakly supervised methods and even is comparable with some fully supervised approaches. Our source code and datasets are available at: https://github.com/shuyonggao/PVSOD.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- ZOOM: Learning Video Mirror Detection with Extremely-Weak SupervisionKe Xu, Tsun Wai Siu, Rynson W. H. LauAAAI 2024 · 被引用 10 次
- Scoring, Remember, and Reference: Catching Camouflaged Objects in VideosYu'ang Feng, Shuyong Gao, Fuzhen Yan, Yicheng Song 等ICCV 2025 · 被引用 2 次
- Bridging RGB and Hematoxylin Components: An Interleaved Guidance and Fusion Framework for Point Supervised Nuclei SegmentationZihan Huan, Xipeng Pan, Hualong Zhang, Siyang Feng 等CVPR 2026
它引用的顶会 Paper23
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan 等ICCV 2021 · 被引用 4,909 次
- Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNetLi Yuan, Yunpeng Chen, Tao Wang, Weihao Yu 等ICCV 2021 · 被引用 2,462 次
相关 Paper
- Weakly Supervised Video Salient Object DetectionWangbo Zhao, Jing Zhang, Long Li, Nick Barnes 等CVPR 2021
- Weakly-Supervised Salient Object Detection Using Point SupervisonShuyong Gao, Wei Zhang, Yan Wang, Qianyu Guo 等AAAI 2022 · 被引用 77 次
- Semi-Supervised Video Salient Object Detection Using Pseudo-LabelsPengxiang Yan, Guanbin Li, Yuan Xie, Zhen Li 等ICCV 2019 · 被引用 134 次
- Motion Guided Attention for Video Salient Object DetectionHaofeng Li, Guanqi Chen, Guanbin Li, Yizhou YuICCV 2019 · 被引用 200 次
- Weakly-Supervised Salient Object Detection via Scribble AnnotationsJing Zhang, Xin Yu, Aixuan Li, Peipei Song 等CVPR 2020
