Video Instance Segmentation Tracking With a Modified VAE Architecture
Chung-Ching Lin, Ying Hung, Rogério Feris, Linglin He
Abstract
We propose a modified variational autoencoder (VAE) architecture built on top of Mask R-CNN for instance-level video segmentation and tracking. The method builds a shared encoder and three parallel decoders, yielding three disjoint branches for predictions of future frames, object detection boxes, and instance segmentation masks. To effectively solve multiple learning tasks, we introduce a Gaussian Process model to enhance the statistical representation of VAE by relaxing the prior strong independent and identically distributed (iid) assumption of conventional VAEs and allowing potential correlations among extracted latent variables. The network learns embedded spatial interdependence and motion continuity in video data and creates a representation that is effective to produce high-quality segmentation masks and track multiple instances in diverse and unstructured videos. Evaluation on a variety of recently in- troduced datasets shows that our model outperforms previous methods and achieves the new best in class performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3443308d-5cd0-4760-8be3-35e9d90179e0Cited by top-tier papers11
- TF-Blender: Temporal Feature Blender for Video Object DetectionYiming Cui, Liqi Yan, Zhiwen Cao, Dongfang LiuICCV 2021 · 171 citations
- A Generalist Framework for Panoptic Segmentation of Images and VideosTing Chen, Lala Li, Saurabh Saxena, Geoffrey E. Hinton et al.ICCV 2023 · 140 citations
- Crossover Learning for Fast Online Video Instance SegmentationShusheng Yang, Yuxin Fang, Xinggang Wang, Yu Li et al.ICCV 2021 · 124 citations
- Video Instance Segmentation with a Propose-Reduce ParadigmHuaijia Lin, Ruizheng Wu, Shu Liu, Jiangbo Lu et al.ICCV 2021 · 110 citations
- Prototypical Cross-Attention Networks for Multiple Object Tracking and SegmentationLei Ke, Xia Li, Martin Danelljan, Yu-Wing Tai et al.NeurIPS 2021 · 92 citations
Builds on1
Related papers
- Classifying, Segmenting, and Tracking Object Instances in Video with Mask PropagationGedas Bertasius, Lorenzo TorresaniCVPR 2020
- End-to-End Video Instance Segmentation via Spatial-Temporal Graph Neural NetworksTao Wang, Ning Xu, Kean Chen, Weiyao LinICCV 2021 · 30 citations
- VONet: Unsupervised Video Object Learning With Parallel U-Net Attention and Object-wise Sequential VAEHaonan Yu, Wei XuICLR 2024 · 1 citation
- Recurrent Video Masked AutoencodersDaniel Zoran, Nikhil Parthasarathy, Yi Yang, Drew A. Hudson et al.CVPR 2026 · 9 citations
- Video Autoencoder: self-supervised disentanglement of static 3D structure and motionZihang Lai, Sifei Liu, Alexei A. Efros, Xiaolong WangICCV 2021 · 37 citations
