Motion-Attentive Transition for Zero-Shot Video Object Segmentation
Tianfei Zhou, Shunzhou Wang, Yi Zhou, Yazhou Yao, Jianwu Li, Ling Shao
摘要
In this paper, we present a novel end-to-end learning neural network, i.e., MATNet, for zero-shot video object segmentation (ZVOS). Motivated by the human visual attention behavior, MATNet leverages motion cues as a bottom-up signal to guide the perception of object appearance. To achieve this, an asymmetric attention block, named Motion-Attentive Transition (MAT), is proposed within a two-stream encoder network to firstly identify moving regions and then attend appearance learning to capture the full extent of objects. Putting MATs in different convolutional layers, our encoder becomes deeply interleaved, allowing for close hierarchical interactions between object apperance and motion. Such a biologically-inspired design is proven to be superb to conventional two-stream structures, which treat motion and appearance independently in separate streams and often suffer severe overfitting to object appearance. Moreover, we introduce a bridge network to modulate multi-scale spatiotemporal features into more compact, discriminative and scale-sensitive representations, which are subsequently fed into a boundary-aware decoder network to produce accurate segmentation with crisp boundaries. We perform extensive quantitative and qualitative experiments on four challenging public benchmarks, i.e., DAVIS16, DAVIS17, FBMS and YouTube-Objects. Results show that our method achieves compelling performance against current state-of-the-art ZVOS methods. To further demonstrate the generalization ability of our spatiotemporal learning framework, we extend MATNet to another relevant task: dynamic visual attention prediction (DVAP). The experiments on two popular datasets (i.e., Hollywood-2 and UCF-Sports) further verify the superiority of our model (our code is available at https://github.com/tfzhou/MATNet).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper37
- MOSE: A New Dataset for Video Object Segmentation in Complex ScenesHenghui Ding, Chang Liu, Shuting He, Xudong Jiang 等ICCV 2023 · 被引用 267 次
- Regional Semantic Contrast and Aggregation for Weakly Supervised Semantic SegmentationTianfei Zhou, Meijie Zhang, Fang Zhao, Jianwu LiCVPR 2022 · 被引用 190 次
- Self-supervised Video Object Segmentation by Motion GroupingCharig Yang, Hala Lamdouar, Erika Lu, Andrew Zisserman 等ICCV 2021 · 被引用 188 次
- Full-Duplex Strategy for Video Object SegmentationGe-Peng Ji, Keren Fu, Zhe Wu, Deng-Ping Fan 等ICCV 2021 · 被引用 173 次
- Group-Wise Semantic Mining for Weakly Supervised Semantic SegmentationXueyi Li, Tianfei Zhou, Jianwu Li, Yi Zhou 等AAAI 2021 · 被引用 143 次
它引用的顶会 Paper1
相关 Paper
- Multi-Source Fusion and Automatic Predictor Selection for Zero-Shot Video Object SegmentationXiaoqi Zhao, Youwei Pang, Jiaxing Yang, Lihe Zhang 等ACM MM 2021 · 被引用 35 次
- Learning Motion-Appearance Co-Attention for Zero-Shot Video Object SegmentationShu Yang, Lu Zhang, Jinqing Qi, Huchuan Lu 等ICCV 2021 · 被引用 76 次
- SSTVOS: Sparse Spatiotemporal Transformers for Video Object SegmentationBrendan Duke, Abdalla Ahmed, Christian Wolf, Parham Aarabi 等CVPR 2021
- F2Net: Learning to Focus on the Foreground for Unsupervised Video Object SegmentationDaizong Liu, Dongdong Yu, Changhu Wang, Pan ZhouAAAI 2021 · 被引用 53 次
- RANet: Ranking Attention Network for Fast Video Object SegmentationZiqin Wang, Jun Xu, Li Liu, Fan Zhu 等ICCV 2019 · 被引用 217 次
