Target-Aware Object Discovery and Association for Unsupervised Video Multi-Object Segmentation
Tianfei Zhou, Jianwu Li, Xueyi Li, Ling Shao
摘要
This paper addresses the task of unsupervised video multi-object segmentation. Current approaches follow a two-stage paradigm: 1) detect object proposals using pre-trained Mask R-CNN, and 2) conduct generic feature matching for temporal association using re-identification techniques. However, the generic features, widely used in both stages, are not reliable for characterizing unseen objects, leading to poor generalization. To address this, we introduce a novel approach for more accurate and efficient spatio-temporal segmentation. In particular, to address instance discrimination, we propose to combine foreground region estimation and instance grouping together in one network, and additionally introduce temporal guidance for segmenting each frame, enabling more accurate object discovery. For temporal association, we complement current video object segmentation architectures with a discriminative appearance model, capable of capturing more finegrained target-specific information. Given object proposals from the instance discrimination network, three essential strategies are adopted to achieve accurate segmentation: 1) target-specific tracking using a memory-augmented appearance model; 2) target-agnostic verification to trace possible tracklets for the proposal; 3) adaptive memory updating using the verified segments. We evaluate the proposed approach on DAVIS 17 and YouTube-VIS, and the results demonstrate that it outperforms state-of-the-art methods both in segmentation accuracy and inference speed.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Regional Semantic Contrast and Aggregation for Weakly Supervised Semantic SegmentationTianfei Zhou, Meijie Zhang, Fang Zhao, Jianwu LiCVPR 2022 · 被引用 190 次
- Locality-Aware Inter-and Intra-Video Reconstruction for Self-Supervised Correspondence LearningLiulei Li, Tianfei Zhou, Wenguan Wang, Lu Yang 等CVPR 2022 · 被引用 41 次
- From ViT Features to Training-free Video Object Segmentation via Streaming-data Mixture ModelsRoy Uziel, Or Dinari, Oren FreifeldNeurIPS 2023 · 被引用 6 次
- Video Object of Interest SegmentationSiyuan Zhou, Chunru Zhan, Biao Wang, Tiezheng Ge 等AAAI 2023
它引用的顶会 Paper15
- Video Instance SegmentationLinjie Yang, Yuchen Fan, Ning XuICCV 2019 · 被引用 615 次
- Exploring Cross-Image Pixel Contrast for Semantic SegmentationWenguan Wang, Tianfei Zhou, Fisher Yu, Jifeng Dai 等ICCV 2021 · 被引用 568 次
- Zero-Shot Video Object Segmentation via Attentive Graph Neural NetworksWenguan Wang, Xiankai Lu, Jianbing Shen, David J. Crandall 等ICCV 2019 · 被引用 294 次
- SSAP: Single-Shot Instance Segmentation With Affinity PyramidNaiyu Gao, Yanhu Shan, Yupei Wang, Xin Zhao 等ICCV 2019 · 被引用 246 次
- Motion-Attentive Transition for Zero-Shot Video Object SegmentationTianfei Zhou, Shunzhou Wang, Yi Zhou, Yazhou Yao 等AAAI 2020 · 被引用 210 次
相关 Paper
- Look Before You Match: Instance Understanding Matters in Video Object SegmentationJunke Wang, Dongdong Chen, Zuxuan Wu, Chong Luo 等CVPR 2023
- Video Object Segmentation with Dynamic Memory Networks and Adaptive Object AlignmentShuxian Liang, Xu Shen, Jianqiang Huang, Xian-Sheng HuaICCV 2021 · 被引用 28 次
- Prototypical Cross-Attention Networks for Multiple Object Tracking and SegmentationLei Ke, Xia Li, Martin Danelljan, Yu-Wing Tai 等NeurIPS 2021 · 被引用 92 次
- Classifying, Segmenting, and Tracking Object Instances in Video with Mask PropagationGedas Bertasius, Lorenzo TorresaniCVPR 2020
- Query-Memory Re-Aggregation for Weakly-supervised Video Object SegmentationFanchao Lin, Hongtao Xie, Yan Li, Yongdong ZhangAAAI 2021 · 被引用 26 次
