Multi-Source Fusion and Automatic Predictor Selection for Zero-Shot Video Object Segmentation
Xiaoqi Zhao, Youwei Pang, Jiaxing Yang, Lihe Zhang, Huchuan Lu
Abstract
Location and appearance are the key cues for video object segmentation. Many sources such as RGB, depth, optical flow and static saliency can provide useful information about the objects. However, existing approaches only utilize the RGB or RGB and optical flow. In this paper, we propose a novel multi-source fusion network for zero-shot video object segmentation. With the help of interoceptive spatial attention module (ISAM), spatial importance of each source is highlighted. Furthermore, we design a feature purification module (FPM) to filter the inter-source incompatible features. By the ISAM and FPM, the multi-source features are effectively fused. In addition, we put forward an automatic predictor selection network (APS) to select the better prediction of either the static saliency predictor or the moving object predictor in order to prevent over-reliance on the failed results caused by low-quality optical flow maps. Extensive experiments on three challenging public benchmarks (i.e. DAVIS, Youtube-Objects and FBMS) show that the proposed model achieves compelling performance against the state-of-the-arts. The source code will be publicly available at https://github.com/Xiaoqi-Zhao-DLUT/Multi-Source-APS-ZVOS
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2ead3472-1380-4209-ac11-2329f42a2a5bCited by top-tier papers9
- Zoom In and Out: A Mixed-scale Triplet Network for Camouflaged Object DetectionYouwei Pang, Xiaoqi Zhao, Tian-Zhu Xiang, Lihe Zhang et al.CVPR 2022 · 417 citations
- Self-Supervised Pretraining for RGB-D Salient Object DetectionXiaoqi Zhao, Youwei Pang, Lihe Zhang, Huchuan Lu et al.AAAI 2022 · 78 citations
- Joint Semantic Mining for Weakly Supervised RGB-D Salient Object DetectionJingjing Li, Wei Ji, Qi Bi, Cheng Yan et al.NeurIPS 2021 · 56 citations
- You Only Infer Once: Cross-Modal Meta-Transfer for Referring Video Object SegmentationDezhuang Li, Ruoqi Li, Lijun Wang, Yifan Wang et al.AAAI 2022 · 54 citations
- Promoting Saliency From Depth: Deep Unsupervised RGB-D Saliency DetectionWei Ji, Jingjing Li, Qi Bi, Chuan Guo et al.ICLR 2022 · 46 citations
Builds on8
- Depth-Induced Multi-Scale Recurrent Attention Network for Saliency DetectionYongri Piao, Wei Ji, Jingjing Li, Miao Zhang et al.ICCV 2019 · 450 citations
- Zero-Shot Video Object Segmentation via Attentive Graph Neural NetworksWenguan Wang, Xiankai Lu, Jianbing Shen, David J. Crandall et al.ICCV 2019 · 294 citations
- Motion-Attentive Transition for Zero-Shot Video Object SegmentationTianfei Zhou, Shunzhou Wang, Yi Zhou, Yazhou Yao et al.AAAI 2020 · 210 citations
- CDTB: A Color and Depth Visual Object Tracking Dataset and BenchmarkAlan Lukezic, Ugur Kart, Jani Käpylä, Ahmed Durmush et al.ICCV 2019 · 79 citations
- Is Depth Really Necessary for Salient Object Detection?Jiawei Zhao, Yifan Zhao, Jia Li, Xiaowu ChenACM MM 2020 · 71 citations
Related papers
- Learning Motion-Appearance Co-Attention for Zero-Shot Video Object SegmentationShu Yang, Lu Zhang, Jinqing Qi, Huchuan Lu et al.ICCV 2021 · 76 citations
- Motion Guided Attention for Video Salient Object DetectionHaofeng Li, Guanqi Chen, Guanbin Li, Yizhou YuICCV 2019 · 200 citations
- Dual Prototype Attention for Unsupervised Video Object SegmentationSuhwan Cho, Minhyeok Lee, Seunghoon Lee, Dogyoon Lee et al.CVPR 2024
- Full-Duplex Strategy for Video Object SegmentationGe-Peng Ji, Keren Fu, Zhe Wu, Deng-Ping Fan et al.ICCV 2021 · 173 citations
- SimulFlow: Simultaneously Extracting Feature and Identifying Target for Unsupervised Video Object SegmentationLingyi Hong, Wei Zhang, Shuyong Gao, Hong Lu et al.ACM MM 2023 · 14 citations
