Learning Hierarchical Embedding for Video Instance Segmentation
Zheyun Qin, Xiankai Lu, Xiushan Nie, Xiantong Zhen, Yilong Yin
Abstract
In this paper, we address video instance segmentation using a new generative model that learns effective representations of the target and background appearance. We propose to exploit hierarchical structural embedding over spatio-temporal space, which is compact, powerful, and flexible in contrast to current tracking-by-detection methods. Specifically, our model segments and tracks instances across space and time in a single forward pass, which is formulated as hierarchical embedding learning. The model is trained to locate the pixels belonging to specific instances over a video clip. We firstly take advantage of a novel mixing function to better fuse spatio-temporal embeddings. Moreover, we introduce normalizing flows to further improve the robustness of the learned appearance embedding, which theoretically extends conventional generative flows to a factorized conditional scheme. Comprehensive experiments on the video instance segmentation benchmark, i.e., YouTube-VIS, demonstrate the effectiveness of the proposed approach. Furthermore, we evaluate our method on an unsupervised video object segmentation dataset to demonstrate its generalizability.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e1eb6e3c-e841-4d43-8c40-a6fee43d30d6Cited by top-tier papers3
- Efficient Video Instance Segmentation via Tracklet Query and ProposalJialian Wu, Sudhir Yarram, Hui Liang, Tian Lan et al.CVPR 2022 · 33 citations
- Exposing the Self-Supervised Space-Time Correspondence Learning via Graph KernelsZheyun Qin, Xiankai Lu, Xiushan Nie, Yilong Yin et al.AAAI 2023 · 22 citations
- Partitioned Saliency Ranking with Dense Pyramid TransformersChengxiao Sun, Yan Xu, Jialun Pei, Haopeng Fang et al.ACM MM 2023 · 4 citations
Builds on6
- PointFlow: 3D Point Cloud Generation With Continuous Normalizing FlowsGuandao Yang, Xun Huang, Zekun Hao, Ming-Yu Liu et al.ICCV 2019 · 794 citations
- Video Instance SegmentationLinjie Yang, Yuchen Fan, Ning XuICCV 2019 · 615 citations
- Zero-Shot Video Object Segmentation via Attentive Graph Neural NetworksWenguan Wang, Xiankai Lu, Jianbing Shen, David J. Crandall et al.ICCV 2019 · 294 citations
- Video Instance Segmentation Tracking With a Modified VAE ArchitectureChung-Ching Lin, Ying Hung, Rogério Feris, Linglin HeCVPR 2020
- Classifying, Segmenting, and Tracking Object Instances in Video with Mask PropagationGedas Bertasius, Lorenzo TorresaniCVPR 2020
Related papers
- End-to-End Video Instance Segmentation via Spatial-Temporal Graph Neural NetworksTao Wang, Ning Xu, Kean Chen, Weiyao LinICCV 2021 · 30 citations
- Crossover Learning for Fast Online Video Instance SegmentationShusheng Yang, Yuxin Fang, Xinggang Wang, Yu Li et al.ICCV 2021 · 124 citations
- Learning to Track Instances without Video AnnotationsYang Fu, Sifei Liu, Umar Iqbal, Shalini De Mello et al.CVPR 2021
- CompFeat: Comprehensive Feature Aggregation for Video Instance SegmentationYang Fu, Linjie Yang, Ding Liu, Thomas S. Huang et al.AAAI 2021 · 77 citations
- VideoCutLER: Surprisingly Simple Unsupervised Video Instance SegmentationXudong Wang, Ishan Misra, Ziyun Zeng, Rohit Girdhar et al.CVPR 2024
