Motion-Based Generator Model: Unsupervised Disentanglement of Appearance, Trackable and Intrackable Motions in Dynamic Patterns
Jianwen Xie, Ruiqi Gao, Zilong Zheng, Song-Chun Zhu, Ying Nian Wu
Abstract
Dynamic patterns are characterized by complex spatial and motion patterns. Understanding dynamic patterns requires a disentangled representational model that separates the factorial components. A commonly used model for dynamic patterns is the state space model, where the state evolves over time according to a transition model and the state generates the observed image frames according to an emission model. To model the motions explicitly, it is natural for the model to be based on the motions or the displacement fields of the pixels. Thus in the emission model, we let the hidden state generate the displacement field, which warps the trackable component in the previous image frame to generate the next frame while adding a simultaneously emitted residual image to account for the change that cannot be explained by the deformation. The warping of the previous image is about the trackable part of the change of image frame, while the residual image is about the intrackable part of the image. We use a maximum likelihood algorithm to learn the model parameters that iterates between inferring latent noise vectors that drive the transition model and updating the parameters given the inferred latent vectors. Meanwhile we adopt a regularization term to penalize the norms of the residual images to encourage the model to explain the change of image frames by trackable motion. Unlike existing methods on dynamic patterns, we learn our model in unsupervised setting without ground truth displacement fields or optical flows. In addition, our model defines a notion of intrackability by the separation of warped component and residual component in each image frame. We show that our method can synthesize realistic dynamic pattern, and disentangling appearance, trackable and intrackable motions. The learned models can be useful for motion transfer, and it is natural to adopt it to define and measure intrackability of a dynamic pattern.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 76c28071-815d-4858-bdeb-e03c3b3b9619Cited by top-tier papers8
- Latent Image Animator: Learning to Animate Images via Latent Space NavigationYaohui Wang, Di Yang, François Brémond, Antitza DantchevaICLR 2022 · 219 citations
- Blind Image Super-resolution with Elaborate Degradation Modeling on Noise and KernelZongsheng Yue, Qian Zhao, Jianwen Xie, Lei Zhang et al.CVPR 2022 · 76 citations
- Vlogger: Make Your Dream A VlogShaobin Zhuang, Kunchang Li, Xinyuan Chen, Yaohui Wang et al.CVPR 2024 · 23 citations
- VideoWorld 2: Learning Transferable Knowledge from Real-world VideosZhongwei Ren, Yunchao Wei, Xiao Yu, Guixun Luo et al.CVPR 2026 · 9 citations
- A Tale of Two Latent Flows: Learning Latent Space Normalizing Flow with Short-Run Langevin Flow for Approximate InferenceJianwen Xie, Yaxuan Zhu, Yifei Xu, Dingcheng Li et al.AAAI 2023 · 6 citations
Related papers
- VDSM: Unsupervised Video Disentanglement With State-Space Modeling and Deep Mixtures of ExpertsMatthew J. Vowels, Necati Cihan Camgöz, Richard BowdenCVPR 2021
- DiLA: Disentangled Latent Action World ModelsTianqiu Zhang, Muyang Lyu, Yufan Zhang, Fang Fang et al.ICML 2026 · 2 citations
- Bitrate-Controlled Diffusion for Disentangling Motion and Content in VideoXiao Li, Qi Chen, Xiulian Peng, Kai Yu et al.ICCV 2025 · 1 citation
- Hamiltonian Latent Operators for content and motion disentanglement in image sequencesAsif Khan, Amos J. StorkeyNeurIPS 2022 · 5 citations
- Decouple Content and Motion for Conditional Image-to-Video GenerationCuifeng Shen, Yulu Gan, Chen Chen, Xiongwei Zhu et al.AAAI 2024 · 13 citations
