Dancing with Still Images: Video Distillation via Static-Dynamic Disentanglement
Ziyu Wang, Yue Xu, Cewu Lu, Yong-Lu Li
Abstract
Recently, dataset distillation has paved the way towards efficient machine learning, especially for image datasets. However, the distillation for videos, characterized by an exclusive temporal dimension, remains an underexplored domain. In this work, we provide the first systematic study of video distillation and introduce a taxonomy to categorize temporal compression. Our investigation reveals that the temporal information is usually not well learned during distillation, and the temporal dimension of synthetic data contributes little. The observations motivate our unified framework of disentangling the dynamic and static information in the videos. It first distills the videos into still images as static memory and then compensates the dynamic and motion information with a learnable dynamic memory block. Our method achieves state-of-the-art on video datasets at different scales, with a notably smaller memory storage budget. Our code is available at https://github.com/yuz/wan/video.distillation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5891fe71-8809-4b90-817d-2b08c5eefcffCited by top-tier papers7
- Beyond Modality Collapse: Representation Blending for Multimodal Dataset DistillationXin Zhang, Ziruo Zhang, Jiawei Du, Zuozhu Liu et al.NeurIPS 2025 · 9 citations
- Relational Database Distillation: From Structured Tables to Condensed Graph DataXinyi Gao, Jingxi Zhang, Lijian Chen, Tong Chen et al.WWW 2026 · 2 citations
- CovMatch: Cross-Covariance Guided Multimodal Dataset Distillation with Trainable Text EncoderYongmin Lee, Hye Won ChungNeurIPS 2025 · 2 citations
- PRISM: Video Dataset Condensation with Progressive Refinement and Insertion for Sparse MotionJaehyun Choi, Jiwan Hur, Gyojin Han, Jaemyung Yu et al.CVPR 2026 · 2 citations
- Multimodal Dataset Distillation via Phased Teacher ModelsShengbin Guo, Hang Zhao, Senqiao Yang, Chenyang Jiang et al.ICLR 2026 · 1 citation
Builds on16
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 2,927 citations
- VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-TrainingZhan Tong, Yibing Song, Jue Wang, Limin WangNeurIPS 2022 · 2,336 citations
- Dataset Condensation with Gradient MatchingBo Zhao, Konda Reddy Mopuri, Hakan BilenICLR 2021 · 684 citations
- Dataset Condensation with Differentiable Siamese AugmentationBo Zhao, Hakan BilenICML 2021 · 390 citations
- Dataset Distillation with Infinitely Wide Convolutional NetworksTimothy Nguyen, Roman Novak, Lechao Xiao, Jaehoon LeeNeurIPS 2021 · 313 citations
Related papers
- WeClick: Weakly-Supervised Video Semantic Segmentation with Click AnnotationsPeidong Liu, Zibin He, Xiyu Yan, Yong Jiang et al.ACM MM 2021 · 9 citations
- DistInit: Learning Video Representations Without a Single Labeled VideoRohit Girdhar, Du Tran, Lorenzo Torresani, Deva RamananICCV 2019 · 59 citations
- Mix-DANN and Dynamic-Modal-Distillation for Video Domain AdaptationYuehao Yin, Bin Zhu, Jingjing Chen, Lechao Cheng et al.ACM MM 2022 · 7 citations
- Disentangling Spatial and Temporal Learning for Efficient Image-to-Video Transfer LearningZhiwu Qing, Shiwei Zhang, Ziyuan Huang, Yingya Zhang et al.ICCV 2023 · 40 citations
- Distilling Dataset into Neural FieldDonghyeok Shin, HeeSun Bae, Gyuwon Sim, Wanmo Kang et al.ICLR 2025
