Towards Nonlinear Disentanglement in Natural Data with Temporal Sparse Coding
David A. Klindt, Lukas Schott, Yash Sharma, Ivan Ustyuzhaninov, Wieland Brendel, Matthias Bethge, Dylan M. Paiton
Abstract
Disentangling the underlying generative factors from data has so far been limited to carefully constructed scenarios. We propose a path towards natural data by first showing that the statistics of natural data provide enough structure to enable disentanglement, both theoretically and empirically. Specifically, we provide evidence that objects in natural movies undergo transitions that are typically small in magnitude with occasional large jumps, which is characteristic of a temporally sparse distribution. Leveraging this finding we provide a novel proof that relies on a sparse prior on temporally adjacent observations to recover the true latent variables up to permutations and sign flips, providing a stronger result than previous work. We show that equipping practical estimation methods with our prior often surpasses the current state-of-the-art on several established benchmark datasets without any impractical assumptions, such as knowledge of the number of changing generative factors. Furthermore, we contribute two new benchmarks, Natural Sprites and KITTI Masks, which integrate the measured natural dynamics to enable disentanglement evaluation with more realistic datasets. We test our theory on these benchmarks and demonstrate improved performance. We also identify non-obvious challenges for current methods in scaling to more natural domains. Taken together our work addresses key issues in disentanglement research for moving towards more natural settings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 278d9b07-6380-4d5d-ae1b-a01a371fa478Cited by top-tier papers80
- Self-Supervised Learning with Data Augmentations Provably Isolates Content from StyleJulius von Kügelgen, Yash Sharma, Luigi Gresele, Wieland Brendel et al.NeurIPS 2021 · 421 citations
- Contrastive Learning Inverts the Data Generating ProcessRoland S. Zimmermann, Yash Sharma, Steffen Schneider, Matthias Bethge et al.ICML 2021 · 264 citations
- Interventional Causal Representation LearningKartik Ahuja, Divyat Mahajan, Yixin Wang, Yoshua BengioICML 2023 · 143 citations
- CITRIS: Causal Identifiability from Temporal Intervened SequencesPhillip Lippe, Sara Magliacane, Sindy Löwe, Yuki M. Asano et al.ICML 2022 · 136 citations
- Nonparametric Identifiability of Causal Representations from Unknown InterventionsJulius von Kügelgen, Michel Besserve, Wendong Liang, Luigi Gresele et al.NeurIPS 2023 · 127 citations
Builds on7
- Video Instance SegmentationLinjie Yang, Yuchen Fan, Ning XuICCV 2019 · 615 citations
- Weakly Supervised Disentanglement with GuaranteesRui Shu, Yining Chen, Abhishek Kumar, Stefano Ermon et al.ICLR 2020 · 148 citations
- ICE-BeeM: Identifiable Conditional Energy-Based Deep Models Based on Nonlinear ICAIlyes Khemakhem, Ricardo Pio Monti, Diederik P. Kingma, Aapo HyvärinenNeurIPS 2020 · 141 citations
- Disentanglement by Nonlinear ICA with General Incompressible-flow Networks (GIN)Peter Sorrenson, Carsten Rother, Ullrich KötheICLR 2020 · 132 citations
- Unsupervised Model Selection for Variational Disentangled Representation LearningSunny Duan, Loic Matthey, Andre Saraiva, Nick Watters et al.ICLR 2020 · 87 citations
Related papers
- Bitrate-Controlled Diffusion for Disentangling Motion and Content in VideoXiao Li, Qi Chen, Xiulian Peng, Kai Yu et al.ICCV 2025 · 1 citation
- Deformable Sprites for Unsupervised Video DecompositionVickie Ye, Zhengqi Li, Richard Tucker, Angjoo Kanazawa et al.CVPR 2022 · 45 citations
- Motion Representations for Articulated AnimationAliaksandr Siarohin, Oliver J. Woodford, Jian Ren, Menglei Chai et al.CVPR 2021
- PhaseMP: Robust 3D Pose Estimation via Phase-conditioned Human Motion PriorMingyi Shi, Sebastian Starke, Yuting Ye, Taku Komura et al.ICCV 2023 · 27 citations
- VDSM: Unsupervised Video Disentanglement With State-Space Modeling and Deep Mixtures of ExpertsMatthew J. Vowels, Necati Cihan Camgöz, Richard BowdenCVPR 2021
