Spatiotemporal Joint Filter Decomposition in 3D Convolutional Neural Networks
Zichen Miao, Ze Wang, Xiuyuan Cheng, Qiang Qiu
Abstract
In this paper, we introduce spatiotemporal joint filter decomposition to decouple spatial and temporal learning, while preserving spatiotemporal dependency in a video. A 3D convolutional filter is now jointly decomposed over a set of spatial and temporal filter atoms respectively. In this way, a 3D convolutional layer becomes three: a temporal atom layer, a spatial atom layer, and a joint coefficient layer, all three remaining convolutional. One obvious arithmetic manipulation allowed in our joint decomposition is to swap spatial or temporal atoms with a set of atoms that have the same number but different sizes, while keeping the remaining unchanged. For example, as shown later, we can now achieve tempo-invariance by simply dilating temporal atoms only. To illustrate this useful atom-swapping property, we further demonstrate how such a decomposition permits the direct learning of 3D CNNs with full-size videos through iterations of two consecutive sub-stages of learning: In the temporal stage, full-temporal downsampled-spatial data are used to learn temporal atoms and joint coefficients while fixing spatial atoms. In the spatial stage, full-spatial downsampled-temporal data are used for spatial atoms and joint coefficients while fixing temporal atoms. We show empirically on multiple action recognition datasets that, the decoupled spatiotemporal learning significantly reduces the model memory footprints, and allows deep 3D CNNs to model high-spatial long-temporal dependency with limited computational resources while delivering comparable performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Generative Quanta Color ImagingVishal Purohit, Junjie Luo, Yiheng Chi, Qi Guo et al.CVPR 2024 · 1 citation
- Sparse Fine-Tuning of Transformers for Generative TasksWei Chen, Jingxi Yu, Zichen Miao, Qiang QiuICCV 2025 · 1 citation
- Tuning Timestep-Distilled Diffusion Model Using Pairwise Sample OptimizationZichen Miao, Zhengyuan Yang, Kevin Lin, Ze Wang et al.ICLR 2025
- Large Convolutional Model Tuning via Filter SubspaceWei Chen, Zichen Miao, Qiang QiuICLR 2025
- Coeff-Tuning: A Graph Filter Subspace View for Tuning Attention-Based Large ModelsZichen Miao, Wei Chen, Qiang QiuCVPR 2025
Builds on9
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 2,049 citations
- Temporal Interlacing NetworkHao Shao, Shengju Qian, Yu LiuAAAI 2020 · 116 citations
- Adaptive Convolutions with Per-pixel Dynamic Filter AtomZe Wang, Zichen Miao, Jun Hu, Qiang QiuICCV 2021 · 22 citations
- Stochastic Conditional Generative Networks with Basis DecompositionZe Wang, Xiuyuan Cheng, Guillermo Sapiro, Qiang QiuICLR 2020 · 19 citations
Related papers
- Grouped Spatial-Temporal Aggregation for Efficient Action RecognitionChenxu Luo, Alan L. YuilleICCV 2019 · 170 citations
- Gate-Shift Networks for Video Action RecognitionSwathikiran Sudhakaran, Sergio Escalera, Oswald LanzCVPR 2020
- Video-to-Image Casting: A Flatting Method for Video AnalysisXu Chen, Chenqiang Gao, Feng Yang, Xiaohan Wang et al.ACM MM 2021 · 3 citations
- Continual Learning with Filter Atom SwappingZichen Miao, Ze Wang, Wei Chen, Qiang QiuICLR 2022 · 35 citations
- SSAN: Separable Self-Attention Network for Video Representation LearningXudong Guo, Xun Guo, Yan LuCVPR 2021
