Deformable Sprites for Unsupervised Video Decomposition
Vickie Ye, Zhengqi Li, Richard Tucker, Angjoo Kanazawa, Noah Snavely
Abstract
We describe a method to extract persistent elements of a dynamic scene from an input video. We represent each scene element as a Deformable Sprite consisting of three components: 1) a 2D texture image for the entire video, 2) per-frame masks for the element, and 3) non-rigid deformations that map the texture image into each video frame. The resulting decomposition allows for applications such as consistent video editing. Deformable Sprites are a type of video auto-encoder model that is optimized on individual videos, and does not require training on a large dataset, nor does it rely on pretrained models. Moreover, our method does not require object masks or other user input, and discovers moving objects of a wider variety than previous work. We evaluate our approach on standard video datasets and show qualitative results on a diverse array of Internet videos.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers35
- Tracking Everything Everywhere All at OnceQianqian Wang, Yen-Yu Chang, Ruojin Cai, Zhengqi Li et al.ICCV 2023 · 238 citations
- GFlow: Recovering 4D World from Monocular VideoShizun Wang, Xingyi Yang, Qiuhong Shen, Zhenxiang Jiang et al.AAAI 2025 · 47 citations
- CoDeF: Content Deformation Fields for Temporally Consistent Video ProcessingHao Ouyang, Qiuyu Wang, Yuxi Xiao, Qingyan Bai et al.CVPR 2024 · 43 citations
- Splatter a Video: Video Gaussian Representation for Versatile ProcessingYang-Tian Sun, Yihua Huang, Lin Ma, Xiaoyang Lyu et al.NeurIPS 2024 · 41 citations
- SpatialTracker: Tracking Any 2D Pixels in 3D SpaceYuxi Xiao, Qianqian Wang, Shangzhan Zhang, Nan Xue et al.CVPR 2024 · 40 citations
Builds on5
- Object-Centric Learning with Slot AttentionFrancesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran et al.NeurIPS 2020 · 1,275 citations
- Self-supervised Video Object Segmentation by Motion GroupingCharig Yang, Hala Lamdouar, Erika Lu, Andrew Zisserman et al.ICCV 2021 · 188 citations
- MarioNette: Self-Supervised Sprite LearningDmitriy Smirnov, Michaël Gharbi, Matthew Fisher, Vitor Guizilini et al.NeurIPS 2021 · 47 citations
- Omnimatte: Associating Objects and Their Effects in VideoErika Lu, Forrester Cole, Tali Dekel, Andrew Zisserman et al.CVPR 2021
- DyStaB: Unsupervised Object Segmentation via Dynamic-Static BootstrappingYanchao Yang, Brian Lai, Stefano SoattoCVPR 2021
Related papers
- Video Autoencoder: self-supervised disentanglement of static 3D structure and motionZihang Lai, Sifei Liu, Alexei A. Efros, Xiaolong WangICCV 2021 · 37 citations
- Unsupervised Layered Image Decomposition into Object PrototypesTom Monnier, Elliot Vincent, Jean Ponce, Mathieu AubryICCV 2021 · 64 citations
- Bitrate-Controlled Diffusion for Disentangling Motion and Content in VideoXiao Li, Qi Chen, Xiulian Peng, Kai Yu et al.ICCV 2025 · 1 citation
- Unsupervised Volumetric AnimationAliaksandr Siarohin, Willi Menapace, Ivan Skorokhodov, Kyle Olszewski et al.CVPR 2023
- Autodecoding Latent 3D Diffusion ModelsEvangelos Ntavelis, Aliaksandr Siarohin, Kyle Olszewski, Chaoyang Wang et al.NeurIPS 2023 · 65 citations
