SRNet: Spatial Relation Network for Efficient Single-stage Instance Segmentation in Videos
Xiaowen Ying, Xin Li, Mooi Choo Chuah
Abstract
The task of instance segmentation in videos aims to consistently identify objects at pixel level throughout the entire video sequence. Existing state-of-the-art methods either follow the tracking-by-detection paradigm to employ multi-stage pipelines or directly train a complex deep model to process the entire video clips as 3D volumes. However, these methods are typically slow and resource-consuming such that they are often limited to offline processing. In this paper, we propose SRNet, a simple and efficient framework for joint segmentation and tracking of object instances in videos. The key to achieving both high efficiency and accuracy in our framework is to formulate the instance segmentation and tracking problem into a unified spatial-relation learning task where each pixel in the current frame relates to its object center, and each object center relates to its location in the previous frame. This unified learning framework allows our framework to perform join instance segmentation and tracking through a single stage while maintaining low overheads among different learning tasks. Our proposed framework can handle two different task settings and demonstrates comparable performance with state-of-the-art methods on two different benchmarks while running significantly faster.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 5feee992-540e-4987-b78e-09715bf6c010Cited by top-tier papers3
- READ: Large-Scale Neural Scene Rendering for Autonomous DrivingZhuopeng Li, Lu Li, Jianke ZhuAAAI 2023 · 78 citations
- Efficient Video Instance Segmentation via Tracklet Query and ProposalJialian Wu, Sudhir Yarram, Hui Liang, Tian Lan et al.CVPR 2022 · 33 citations
- Leveraging GAN Priors for Few-Shot Part SegmentationMengya Han, Heliang Zheng, Chaoyue Wang, Yong Luo et al.ACM MM 2022 · 5 citations
Related papers
- TrackIME: Enhanced Video Point Tracking via Instance Motion EstimationSeong Hyeon Park, Huiwon Jang, Byungwoo Jeon, Sukmin Yun et al.NeurIPS 2024 · 1 citation
- Video Instance SegmentationLinjie Yang, Yuchen Fan, Ning XuICCV 2019 · 615 citations
- End-to-End Video Instance Segmentation via Spatial-Temporal Graph Neural NetworksTao Wang, Ning Xu, Kean Chen, Weiyao LinICCV 2021 · 30 citations
- Classifying, Segmenting, and Tracking Object Instances in Video with Mask PropagationGedas Bertasius, Lorenzo TorresaniCVPR 2020
- SG-Net: Spatial Granularity Network for One-Stage Video Instance SegmentationDongfang Liu, Yiming Cui, Wenbo Tan, Yingjie Victor ChenCVPR 2021
