SC2025Top-tier venue
Generative Latent Diffusion for Efficient Spatiotemporal Data Reduction
Xiao Li, Liangji Zhu, Anand Rangarajan, Sanjay Ranka
Abstract
Generative models have demonstrated strong performance in conditional settings and can be viewed as a form of data compression, where the condition serves as a compact representation. However, their limited controllability and reconstruction accuracy restrict their practical application to data compression. In this work, we propose an efficient latent diffusion framework that bridges this gap by combining a variational autoencoder with a conditional diffusion model. Our method compresses only a small number of keyframes into latent space and uses them as conditioning inputs to reconstruct the remaining frames via generative interpolation, eliminating the need to store latent representations for every frame. This approach enables accurate spatiotemporal reconstruction while significantly reducing storage costs. Experimental results across multiple datasets show that our method achieves up to 10× higher compression ratios than rule-based state-of-the-art compressors such as SZ3, and up to 63% better performance than leading learning-based methods under the same reconstruction error.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext acaa46e4-3956-4396-8ae1-3f2f0db783dfBuilds on8
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan et al.NeurIPS 2022 · 2,948 citations
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun et al.ICCV 2021 · 2,947 citations
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 2,927 citations
- MCVD - Masked Conditional Video Diffusion for Prediction, Generation, and InterpolationVikram Voleti, Alexia Jolicoeur-Martineau, Chris PalNeurIPS 2022 · 434 citations
Related papers
- REGEN: Learning Compact Video Embedding with (Re-)Generative DecoderYitian Zhang, Long Mai, Aniruddha Mahapatra, David Bourgin et al.ICCV 2025
- Progressive Growing of Video Tokenizers for Temporally Compact Latent SpacesAniruddha Mahapatra, Long Mai, David Bourgin, Yitian Zhang et al.ICCV 2025 · 3 citations
- CV-VAE: A Compatible Video VAE for Latent Generative Video ModelsSijie Zhao, Yong Zhang, Xiaodong Cun, Shaoshu Yang et al.NeurIPS 2024 · 82 citations
- Realtime Video Frame Interpolation using One-Step Diffusion SamplingYongrui Ma, Shijie Zhao, Mingde Yao, Junlin Li et al.ICLR 2026
- Lost in Latent Space: An Empirical Study of Latent Diffusion Models for Physics EmulationFrançois Rozet, Ruben Ohana, Michael McCabe, Gilles Louppe et al.NeurIPS 2025 · 23 citations
