S3VAE: Self-Supervised Sequential VAE for Representation Disentanglement and Data Generation
Yizhe Zhu, Martin Renqiang Min, Asim Kadav, Hans Peter Graf
摘要
We propose a sequential variational autoencoder to learn disentangled representations of sequential data (e.g., videos and audios) under self-supervision. Specifically, we exploit the benefits of some readily accessible supervisory signals from input data itself or some off-the-shelf functional models and accordingly design auxiliary tasks for our model to utilize these signals. With the supervision of the signals, our model can easily disentangle the representation of an input sequence into static factors and dynamic factors (i.e., time-invariant and time-varying parts). Comprehensive experiments across videos and audios verify the effectiveness of our model on representation disentanglement and generation of sequential data, and demonstrate that, our model with self-supervision performs comparable to, if not better than, the fully-supervised model with ground truth labels, and outperforms state-of-the-art unsupervised models by a large margin.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper31
- Audio2Gestures: Generating Diverse Gestures from Speech Audio with Conditional Variational AutoencodersJing Li, Di Kang, Wenjie Pei, Xuefei Zhe 等ICCV 2021 · 被引用 144 次
- SIMONe: View-Invariant, Temporally-Abstracted Object Representations via Unsupervised Video DecompositionRishabh Kabra, Daniel Zoran, Goker Erdogan, Loic Matthey 等NeurIPS 2021 · 被引用 90 次
- Self-Supervised Learning Disentangled Group Representation as FeatureTan Wang, Zhongqi Yue, Jianqiang Huang, Qianru Sun 等NeurIPS 2021 · 被引用 78 次
- Contrastively Disentangled Sequential Variational AutoencoderJunwen Bai, Weiran Wang, Carla P. GomesNeurIPS 2021 · 被引用 60 次
- Unsupervised Video Domain Adaptation for Action Recognition: A Disentanglement PerspectivePengfei Wei, Lingdong Kong, Xinghua Qu, Yi Ren 等NeurIPS 2023 · 被引用 39 次
它引用的顶会 Paper1
相关 Paper
- Disentangled Recurrent Wasserstein AutoencoderJun Han, Martin Renqiang Min, Ligong Han, Li Erran Li 等ICLR 2021 · 被引用 37 次
- Sample and Predict Your Latent: Modality-free Sequential Disentanglement via Contrastive EstimationIlan Naiman, Nimrod Berman, Omri AzencotICML 2023 · 被引用 12 次
- Sequential Disentanglement by Extracting Static Information From A Single Sequence ElementNimrod Berman, Ilan Naiman, Idan Arbiv, Gal Fadlon 等ICML 2024 · 被引用 9 次
- Video Autoencoder: self-supervised disentanglement of static 3D structure and motionZihang Lai, Sifei Liu, Alexei A. Efros, Xiaolong WangICCV 2021 · 被引用 37 次
- DiffSDA: Unsupervised Diffusion Sequential Disentanglement Across ModalitiesHedi Zisling, Ilan Naiman, Nimrod Berman, Supasorn Suwajanakorn 等ICLR 2026 · 被引用 2 次
