Active Patterns Perceived for Stochastic Video Prediction
Yechao Xu, Zhengxing Sun, Qian Li, Yunhan Sun, Shoutong Luo
摘要
Predicting future scenes based on historical frames is challenging, especially when it comes to the complex uncertainty in nature. We observe that there is a divergence between spatial-temporal variations of active patterns and non-active patterns in a video, where these patterns constitute visual content and the former ones implicate more violent movement. This divergence enables active patterns the higher potential to act with more severe future uncertainty. Meanwhile, the existence of non-active patterns provides an opportunity for machines to examine some underlying rules with a mutual constraint between non-active patterns and active patterns. In order to solve this divergence, we provide a method called active patterns-perceived stochastic video prediction (ASVP) which allows active patterns to be perceived by neural networks during training. Our method starts with separating active patterns along with non-active ones from a video. Then, both scene-based prediction and active pattern-perceived prediction are conducted to respectively capture the variations within the whole scene and active patterns. Specially for active pattern-perceived prediction, a conditional generative adversarial network (CGAN) is exploited to model active patterns as conditions, with a variational autoencoder (VAE) for predicting the complex dynamics of active patterns. Additionally, a mutual constraint is designed to improve the learning procedure for the network to better understand underlying interacting rules among these patterns. Extensive experiments are conducted on both KTH human action and BAIR action-free robot pushing datasets with comparison to state-of-the-art works. Experimental results demonstrate the competitive performance of the proposed method as we expected. The released code and models are at https://github.com/tolearnmuch/ASVP.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper1
问问它们各自怎么用它相关 Paper
- SLAMP: Stochastic Latent Appearance and Motion PredictionAdil Kaan Akan, Erkut Erdem, Aykut Erdem, Fatma GüneyICCV 2021 · 被引用 46 次
- Weakly-supervised Action Transition Learning for Stochastic Human Motion PredictionWei Mao, Miaomiao Liu, Mathieu SalzmannCVPR 2022 · 被引用 29 次
- Self-Supervised Video GANs: Learning for Appearance Consistency and Motion CoherencySangeek Hyun, Jihwan Kim, Jae-Pil HeoCVPR 2021
- Playable Video GenerationWilli Menapace, Stéphane Lathuilière, Sergey Tulyakov, Aliaksandr Siarohin 等CVPR 2021
- Predicting the Future: A Jointly Learnt Model for Action AnticipationHarshala Gammulle, Simon Denman, Sridha Sridharan, Clinton FookesICCV 2019 · 被引用 93 次
