Towards Nonlinear Disentanglement in Natural Data with Temporal Sparse Coding
David A. Klindt, Lukas Schott, Yash Sharma, Ivan Ustyuzhaninov, Wieland Brendel, Matthias Bethge, Dylan M. Paiton
摘要
Disentangling the underlying generative factors from data has so far been limited to carefully constructed scenarios. We propose a path towards natural data by first showing that the statistics of natural data provide enough structure to enable disentanglement, both theoretically and empirically. Specifically, we provide evidence that objects in natural movies undergo transitions that are typically small in magnitude with occasional large jumps, which is characteristic of a temporally sparse distribution. Leveraging this finding we provide a novel proof that relies on a sparse prior on temporally adjacent observations to recover the true latent variables up to permutations and sign flips, providing a stronger result than previous work. We show that equipping practical estimation methods with our prior often surpasses the current state-of-the-art on several established benchmark datasets without any impractical assumptions, such as knowledge of the number of changing generative factors. Furthermore, we contribute two new benchmarks, Natural Sprites and KITTI Masks, which integrate the measured natural dynamics to enable disentanglement evaluation with more realistic datasets. We test our theory on these benchmarks and demonstrate improved performance. We also identify non-obvious challenges for current methods in scaling to more natural domains. Taken together our work addresses key issues in disentanglement research for moving towards more natural settings.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper80
- Self-Supervised Learning with Data Augmentations Provably Isolates Content from StyleJulius von Kügelgen, Yash Sharma, Luigi Gresele, Wieland Brendel 等NeurIPS 2021 · 被引用 421 次
- Contrastive Learning Inverts the Data Generating ProcessRoland S. Zimmermann, Yash Sharma, Steffen Schneider, Matthias Bethge 等ICML 2021 · 被引用 264 次
- Interventional Causal Representation LearningKartik Ahuja, Divyat Mahajan, Yixin Wang, Yoshua BengioICML 2023 · 被引用 143 次
- CITRIS: Causal Identifiability from Temporal Intervened SequencesPhillip Lippe, Sara Magliacane, Sindy Löwe, Yuki M. Asano 等ICML 2022 · 被引用 136 次
- Nonparametric Identifiability of Causal Representations from Unknown InterventionsJulius von Kügelgen, Michel Besserve, Wendong Liang, Luigi Gresele 等NeurIPS 2023 · 被引用 127 次
它引用的顶会 Paper7
- Video Instance SegmentationLinjie Yang, Yuchen Fan, Ning XuICCV 2019 · 被引用 615 次
- Weakly Supervised Disentanglement with GuaranteesRui Shu, Yining Chen, Abhishek Kumar, Stefano Ermon 等ICLR 2020 · 被引用 148 次
- ICE-BeeM: Identifiable Conditional Energy-Based Deep Models Based on Nonlinear ICAIlyes Khemakhem, Ricardo Pio Monti, Diederik P. Kingma, Aapo HyvärinenNeurIPS 2020 · 被引用 141 次
- Disentanglement by Nonlinear ICA with General Incompressible-flow Networks (GIN)Peter Sorrenson, Carsten Rother, Ullrich KötheICLR 2020 · 被引用 132 次
- Unsupervised Model Selection for Variational Disentangled Representation LearningSunny Duan, Loic Matthey, Andre Saraiva, Nick Watters 等ICLR 2020 · 被引用 87 次
相关 Paper
- Bitrate-Controlled Diffusion for Disentangling Motion and Content in VideoXiao Li, Qi Chen, Xiulian Peng, Kai Yu 等ICCV 2025 · 被引用 1 次
- Deformable Sprites for Unsupervised Video DecompositionVickie Ye, Zhengqi Li, Richard Tucker, Angjoo Kanazawa 等CVPR 2022 · 被引用 45 次
- Motion Representations for Articulated AnimationAliaksandr Siarohin, Oliver J. Woodford, Jian Ren, Menglei Chai 等CVPR 2021
- PhaseMP: Robust 3D Pose Estimation via Phase-conditioned Human Motion PriorMingyi Shi, Sebastian Starke, Yuting Ye, Taku Komura 等ICCV 2023 · 被引用 27 次
- VDSM: Unsupervised Video Disentanglement With State-Space Modeling and Deep Mixtures of ExpertsMatthew J. Vowels, Necati Cihan Camgöz, Richard BowdenCVPR 2021
