Full-Duplex Strategy for Video Object Segmentation
Ge-Peng Ji, Keren Fu, Zhe Wu, Deng-Ping Fan, Jianbing Shen, Ling Shao
Abstract
Previous video object segmentation approaches mainly focus on using simplex solutions between appearance and motion, limiting feature collaboration efficiency among and across these two cues. In this work, we study a novel and efficient full-duplex strategy network (FSNet) to address this issue, by considering a better mutual restraint scheme between motion and appearance in exploiting the crossmodal features from the fusion and decoding stage. Specifically, we introduce the relational cross-attention module (RCAM) to achieve bidirectional message propagation across embedding sub-spaces. To improve the model's robustness and update the inconsistent features from the spatial-temporal embeddings, we adopt the bidirectional purification module (BPM) after the RCAM. Extensive experiments on five popular benchmarks show that our FSNet is robust to various challenging scenarios (e.g., motion blur, occlusion) and achieves favourable performance against existing cutting-edges both in the video object segmentation and video salient object detection tasks. The project is publicly available at: https://dpfan.net/FSNet .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d2035fb3-0379-465e-866a-e029971fe7efCited by top-tier papers31
- Zoom In and Out: A Mixed-scale Triplet Network for Camouflaged Object DetectionYouwei Pang, Xiaoqi Zhao, Tian-Zhu Xiang, Lihe Zhang et al.CVPR 2022 · 417 citations
- Dynamic Context-Sensitive Filtering Network for Video Salient Object DetectionMiao Zhang, Jie Liu, Yifei Wang, Yongri Piao et al.ICCV 2021 · 112 citations
- Pyramid Grafting Network for One-Stage High Resolution Saliency DetectionChenxi Xie, Changqun Xia, Mingcan Ma, Zhirui Zhao et al.CVPR 2022 · 112 citations
- VSCode: General Visual Salient and Camouflaged Object Detection with 2D Prompt LearningZiyang Luo, Nian Liu, Wangbo Zhao, Xuguang Yang et al.CVPR 2024 · 96 citations
- Implicit Motion Handling for Video Camouflaged Object DetectionXuelian Cheng, Huan Xiong, Deng-Ping Fan, Yiran Zhong et al.CVPR 2022 · 83 citations
Builds on26
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan et al.ICCV 2021 · 4,909 citations
- Video Instance SegmentationLinjie Yang, Yuchen Fan, Ning XuICCV 2019 · 615 citations
- Stacked Cross Refinement Network for Edge-Aware Salient Object DetectionZhe Wu, Li Su, Qingming HuangICCV 2019 · 374 citations
- Zero-Shot Video Object Segmentation via Attentive Graph Neural NetworksWenguan Wang, Xiankai Lu, Jianbing Shen, David J. Crandall et al.ICCV 2019 · 294 citations
- Motion-Attentive Transition for Zero-Shot Video Object SegmentationTianfei Zhou, Shunzhou Wang, Yi Zhou, Yazhou Yao et al.AAAI 2020 · 210 citations
Related papers
- Learning Motion-Appearance Co-Attention for Zero-Shot Video Object SegmentationShu Yang, Lu Zhang, Jinqing Qi, Huchuan Lu et al.ICCV 2021 · 76 citations
- Motion Deblurring via Spatial-Temporal Collaboration of Frames and EventsWen Yang, Jinjian Wu, Jupo Ma, Leida Li et al.AAAI 2024 · 19 citations
- Multi-Source Fusion and Automatic Predictor Selection for Zero-Shot Video Object SegmentationXiaoqi Zhao, Youwei Pang, Jiaxing Yang, Lihe Zhang et al.ACM MM 2021 · 35 citations
- F2Net: Learning to Focus on the Foreground for Unsupervised Video Object SegmentationDaizong Liu, Dongdong Yu, Changhu Wang, Pan ZhouAAAI 2021 · 53 citations
- CCNet: Criss-Cross Attention for Semantic SegmentationZilong Huang, Xinggang Wang, Lichao Huang, Chang Huang et al.ICCV 2019 · 2,972 citations
