Clockwork Variational Autoencoders
Vaibhav Saxena, Jimmy Ba, Danijar Hafner
Abstract
Deep learning has enabled algorithms to generate realistic images. However, accurately predicting long video sequences requires understanding long-term dependencies and remains an open challenge. While existing video prediction models succeed at generating sharp images, they tend to fail at accurately predicting far into the future. We introduce the Clockwork VAE (CW-VAE), a video prediction model that leverages a hierarchy of latent sequences, where higher levels tick at slower intervals. We demonstrate the benefits of both hierarchical latents and temporal abstraction on 4 diverse video prediction datasets with sequences of up to 1000 frames, where CW-VAE outperforms top video prediction models. Additionally, we propose a Minecraft benchmark for long-term video prediction. We conduct several experiments to gain insights into CW-VAE and confirm that slower levels learn to represent objects that change more slowly in the video, and faster levels learn to represent faster objects.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1d7ccf70-f312-4874-a42e-6535c58680ecCited by top-tier papers22
- Flexible Diffusion Modeling of Long VideosWilliam Harvey, Saeid Naderiparizi, Vaden Masrani, Christian Weilbach et al.NeurIPS 2022 · 384 citations
- Discovering and Achieving Goals via World ModelsRussell Mendonca, Oleh Rybkin, Kostas Daniilidis, Danijar Hafner et al.NeurIPS 2021 · 177 citations
- Video Language PlanningYilun Du, Sherry Yang, Pete Florence, Fei Xia et al.ICLR 2024 · 161 citations
- Simple Hierarchical Planning with DiffusionChang Chen, Fei Deng, Kenji Kawaguchi, Caglar Gulcehre et al.ICLR 2024 · 79 citations
- Facing Off World Model Backbones: RNNs, Transformers, and S4Fei Deng, Junyeong Park, Sungjin AhnNeurIPS 2023 · 53 citations
Builds on8
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 1,852 citations
- Mastering Atari with Discrete World ModelsDanijar Hafner, Timothy P. Lillicrap, Mohammad Norouzi, Jimmy BaICLR 2021 · 1,170 citations
- NVAE: A Deep Hierarchical Variational AutoencoderArash Vahdat, Jan KautzNeurIPS 2020 · 1,141 citations
- Scaling Autoregressive Video ModelsDirk Weissenborn, Oscar Täckström, Jakob UszkoreitICLR 2020 · 252 citations
- Improved Conditional VRNNs for Video PredictionLluís Castrejón, Nicolas Ballas, Aaron C. CourvilleICCV 2019 · 177 citations
Related papers
- Revisiting Hierarchical Approach for Persistent Long-Term Video PredictionWonkwang Lee, Whie Jung, Han Zhang, Ting Chen et al.ICLR 2021 · 29 citations
- Temporally Consistent Transformers for Video GenerationWilson Yan, Danijar Hafner, Stephen James, Pieter AbbeelICML 2023 · 47 citations
- Deep Hierarchical Video CompressionMing Lu, Zhihao Duan, Fengqing Zhu, Zhan MaAAAI 2024 · 19 citations
- LV-MAE: Learning Long Video Representations Through Masked-Embedding AutoencodersIlan Naiman, Emanuel Ben Baruch, Oron Anschel, Alon Shoshan et al.ICCV 2025 · 2 citations
- Generating Long Videos of Dynamic ScenesTim Brooks, Janne Hellsten, Miika Aittala, Ting-Chun Wang et al.NeurIPS 2022 · 152 citations
