Video Instance Segmentation Tracking With a Modified VAE Architecture
Chung-Ching Lin, Ying Hung, Rogério Feris, Linglin He
摘要
We propose a modified variational autoencoder (VAE) architecture built on top of Mask R-CNN for instance-level video segmentation and tracking. The method builds a shared encoder and three parallel decoders, yielding three disjoint branches for predictions of future frames, object detection boxes, and instance segmentation masks. To effectively solve multiple learning tasks, we introduce a Gaussian Process model to enhance the statistical representation of VAE by relaxing the prior strong independent and identically distributed (iid) assumption of conventional VAEs and allowing potential correlations among extracted latent variables. The network learns embedded spatial interdependence and motion continuity in video data and creates a representation that is effective to produce high-quality segmentation masks and track multiple instances in diverse and unstructured videos. Evaluation on a variety of recently in- troduced datasets shows that our model outperforms previous methods and achieves the new best in class performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- TF-Blender: Temporal Feature Blender for Video Object DetectionYiming Cui, Liqi Yan, Zhiwen Cao, Dongfang LiuICCV 2021 · 被引用 171 次
- A Generalist Framework for Panoptic Segmentation of Images and VideosTing Chen, Lala Li, Saurabh Saxena, Geoffrey E. Hinton 等ICCV 2023 · 被引用 140 次
- Crossover Learning for Fast Online Video Instance SegmentationShusheng Yang, Yuxin Fang, Xinggang Wang, Yu Li 等ICCV 2021 · 被引用 124 次
- Video Instance Segmentation with a Propose-Reduce ParadigmHuaijia Lin, Ruizheng Wu, Shu Liu, Jiangbo Lu 等ICCV 2021 · 被引用 110 次
- Prototypical Cross-Attention Networks for Multiple Object Tracking and SegmentationLei Ke, Xia Li, Martin Danelljan, Yu-Wing Tai 等NeurIPS 2021 · 被引用 92 次
它引用的顶会 Paper1
相关 Paper
- Classifying, Segmenting, and Tracking Object Instances in Video with Mask PropagationGedas Bertasius, Lorenzo TorresaniCVPR 2020
- End-to-End Video Instance Segmentation via Spatial-Temporal Graph Neural NetworksTao Wang, Ning Xu, Kean Chen, Weiyao LinICCV 2021 · 被引用 30 次
- VONet: Unsupervised Video Object Learning With Parallel U-Net Attention and Object-wise Sequential VAEHaonan Yu, Wei XuICLR 2024 · 被引用 1 次
- Recurrent Video Masked AutoencodersDaniel Zoran, Nikhil Parthasarathy, Yi Yang, Drew A. Hudson 等CVPR 2026 · 被引用 9 次
- Video Autoencoder: self-supervised disentanglement of static 3D structure and motionZihang Lai, Sifei Liu, Alexei A. Efros, Xiaolong WangICCV 2021 · 被引用 37 次
