Clockwork Variational Autoencoders
Vaibhav Saxena, Jimmy Ba, Danijar Hafner
摘要
Deep learning has enabled algorithms to generate realistic images. However, accurately predicting long video sequences requires understanding long-term dependencies and remains an open challenge. While existing video prediction models succeed at generating sharp images, they tend to fail at accurately predicting far into the future. We introduce the Clockwork VAE (CW-VAE), a video prediction model that leverages a hierarchy of latent sequences, where higher levels tick at slower intervals. We demonstrate the benefits of both hierarchical latents and temporal abstraction on 4 diverse video prediction datasets with sequences of up to 1000 frames, where CW-VAE outperforms top video prediction models. Additionally, we propose a Minecraft benchmark for long-term video prediction. We conduct several experiments to gain insights into CW-VAE and confirm that slower levels learn to represent objects that change more slowly in the video, and faster levels learn to represent faster objects.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper22
- Flexible Diffusion Modeling of Long VideosWilliam Harvey, Saeid Naderiparizi, Vaden Masrani, Christian Weilbach 等NeurIPS 2022 · 被引用 384 次
- Discovering and Achieving Goals via World ModelsRussell Mendonca, Oleh Rybkin, Kostas Daniilidis, Danijar Hafner 等NeurIPS 2021 · 被引用 177 次
- Video Language PlanningYilun Du, Sherry Yang, Pete Florence, Fei Xia 等ICLR 2024 · 被引用 161 次
- Simple Hierarchical Planning with DiffusionChang Chen, Fei Deng, Kenji Kawaguchi, Caglar Gulcehre 等ICLR 2024 · 被引用 79 次
- Facing Off World Model Backbones: RNNs, Transformers, and S4Fei Deng, Junyeong Park, Sungjin AhnNeurIPS 2023 · 被引用 53 次
它引用的顶会 Paper8
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 被引用 1,852 次
- Mastering Atari with Discrete World ModelsDanijar Hafner, Timothy P. Lillicrap, Mohammad Norouzi, Jimmy BaICLR 2021 · 被引用 1,170 次
- NVAE: A Deep Hierarchical Variational AutoencoderArash Vahdat, Jan KautzNeurIPS 2020 · 被引用 1,141 次
- Scaling Autoregressive Video ModelsDirk Weissenborn, Oscar Täckström, Jakob UszkoreitICLR 2020 · 被引用 252 次
- Improved Conditional VRNNs for Video PredictionLluís Castrejón, Nicolas Ballas, Aaron C. CourvilleICCV 2019 · 被引用 177 次
相关 Paper
- Revisiting Hierarchical Approach for Persistent Long-Term Video PredictionWonkwang Lee, Whie Jung, Han Zhang, Ting Chen 等ICLR 2021 · 被引用 29 次
- Temporally Consistent Transformers for Video GenerationWilson Yan, Danijar Hafner, Stephen James, Pieter AbbeelICML 2023 · 被引用 47 次
- Deep Hierarchical Video CompressionMing Lu, Zhihao Duan, Fengqing Zhu, Zhan MaAAAI 2024 · 被引用 19 次
- LV-MAE: Learning Long Video Representations Through Masked-Embedding AutoencodersIlan Naiman, Emanuel Ben Baruch, Oron Anschel, Alon Shoshan 等ICCV 2025 · 被引用 2 次
- Generating Long Videos of Dynamic ScenesTim Brooks, Janne Hellsten, Miika Aittala, Ting-Chun Wang 等NeurIPS 2022 · 被引用 152 次
