Sample and Predict Your Latent: Modality-free Sequential Disentanglement via Contrastive Estimation
Ilan Naiman, Nimrod Berman, Omri Azencot
Abstract
Unsupervised disentanglement is a long-standing challenge in representation learning. Recently, self-supervised techniques achieved impressive results in the sequential setting, where data is time-dependent. However, the latter methods employ modality-based data augmentations and random sampling or solve auxiliary tasks. In this work, we propose to avoid that by generating, sampling, and comparing empirical distributions from the underlying variational model. Unlike existing work, we introduce a self-supervised sequential disentanglement framework based on contrastive estimation with no external signals, while using common batch sizes and samples from the latent space itself. In practice, we propose a unified, efficient, and easy-to-code sampling strategy for semantically similar and dissimilar views of the data. We evaluate our approach on video, audio, and time series benchmarks. Our method presents state-of-the-art results in comparison to existing techniques. The code is available at https://github.com/azencot-group/SPYL.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext da1eced4-562c-4536-babf-da7d65ca0309Cited by top-tier papers5
- Sequential Disentanglement by Extracting Static Information From A Single Sequence ElementNimrod Berman, Ilan Naiman, Idan Arbiv, Gal Fadlon et al.ICML 2024 · 9 citations
- First-Order Manifold Data Augmentation for Regression LearningIlya Kaufman, Omri AzencotICML 2024 · 6 citations
- Towards General Modality Translation with Contrastive and Predictive Latent Diffusion BridgeNimrod Berman, Omkar Joglekar, Eitan Kosman, Dotan Di Castro et al.NeurIPS 2025 · 5 citations
- DiffSDA: Unsupervised Diffusion Sequential Disentanglement Across ModalitiesHedi Zisling, Ilan Naiman, Nimrod Berman, Supasorn Suwajanakorn et al.ICLR 2026 · 2 citations
- Bitrate-Controlled Diffusion for Disentangling Motion and Content in VideoXiao Li, Qi Chen, Xiulian Peng, Kai Yu et al.ICCV 2025 · 1 citation
Builds on21
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal et al.NeurIPS 2020 · 5,249 citations
- What Makes for Good Views for Contrastive Learning?Yonglong Tian, Chen Sun, Ben Poole, Dilip Krishnan et al.NeurIPS 2020 · 1,631 citations
- NVAE: A Deep Hierarchical Variational AutoencoderArash Vahdat, Jan KautzNeurIPS 2020 · 1,141 citations
Related papers
- S3VAE: Self-Supervised Sequential VAE for Representation Disentanglement and Data GenerationYizhe Zhu, Martin Renqiang Min, Asim Kadav, Hans Peter GrafCVPR 2020
- Contrastively Disentangled Sequential Variational AutoencoderJunwen Bai, Weiran Wang, Carla P. GomesNeurIPS 2021 · 60 citations
- Spatiotemporal Contrastive Video Representation LearningRui Qian, Tianjian Meng, Boqing Gong, Ming-Hsuan Yang et al.CVPR 2021
- Video Autoencoder: self-supervised disentanglement of static 3D structure and motionZihang Lai, Sifei Liu, Alexei A. Efros, Xiaolong WangICCV 2021 · 37 citations
- Unsupervised Video Domain Adaptation for Action Recognition: A Disentanglement PerspectivePengfei Wei, Lingdong Kong, Xinghua Qu, Yi Ren et al.NeurIPS 2023 · 39 citations
