A Deeper Dive Into What Deep Spatiotemporal Networks Encode: Quantifying Static vs. Dynamic Information
Matthew Kowal, Mennatullah Siam, Md. Amirul Islam, Neil D. B. Bruce, Richard P. Wildes, Konstantinos G. Derpanis
Abstract
Deep spatiotemporal models are used in a variety of computer vision tasks, such as action recognition and video object segmentation. Currently, there is a limited understanding of what information is captured by these models in their intermediate representations. For example, while it has been observed that action recognition algorithms are heavily influenced by visual appearance in single static frames, there is no quantitative methodology for evaluating such static bias in the latent representation compared to bias toward dynamic information (e.g. motion). We tackle this challenge by proposing a novel approach for quantifying the static and dynamic biases of any spatiotemporal model. To show the efficacy of our approach, we analyse two widely studied tasks, action recognition and video object segmentation. Our key findings are threefold: (i) Most examined spatiotemporal models are biased toward static information; although, certain two-stream architectures with cross-connections show a better balance between the static and dynamic information captured. (ii) Some datasets that are commonly assumed to be biased toward dynamics are actually biased toward static information. (iii) Individual units (channels) in an architecture can be biased toward static, dynamic or a combination of the two. <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup> <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup> Project page and code
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers10
- CAST: Cross-Attention in Space and Time for Video Action RecognitionDongho Lee, Jongseo Lee, Jinwoo ChoiNeurIPS 2023 · 43 citations
- Mitigating and Evaluating Static Bias of Action Representations in the Background and the ForegroundHaoxin Li, Yuan Liu, Hanwang Zhang, Boyang LiICCV 2023 · 30 citations
- Opening the Vocabulary of Egocentric ActionsDibyadip Chatterjee, Fadime Sener, Shugao Ma, Angela YaoNeurIPS 2023 · 28 citations
- Dissecting Multimodality in VideoQA Transformer Models by Impairing Modality FusionIshaan Singh Rawal, Alexander Matyasko, Shantanu Jaiswal, Basura Fernando et al.ICML 2024 · 8 citations
- CamSAM2: Segment Anything Accurately in Camouflaged VideosYuli Zhou, Yawei Li, Yuqian Fu, Luca Benini et al.NeurIPS 2025 · 8 citations
Builds on10
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 2,927 citations
- Multiscale Vision TransformersHaoqi Fan, Bo Xiong, Karttikeya Mangalam, Yanghao Li et al.ICCV 2021 · 1,611 citations
- Keeping Your Eye on the Ball: Trajectory Attention in Video TransformersMandela Patrick, Dylan Campbell, Yuki M. Asano, Ishan Misra et al.NeurIPS 2021 · 382 citations
- Motion-Attentive Transition for Zero-Shot Video Object SegmentationTianfei Zhou, Shunzhou Wang, Yi Zhou, Yazhou Yao et al.AAAI 2020 · 210 citations
Related papers
- Time Is MattEr: Temporal Self-supervision for Video TransformersSukmin Yun, Jaehyung Kim, Dongyoon Han, Hwanjun Song et al.ICML 2022 · 17 citations
- Streaming Video ModelYucheng Zhao, Chong Luo, Chuanxin Tang, Dongdong Chen et al.CVPR 2023
- Grouped Spatial-Temporal Aggregation for Efficient Action RecognitionChenxu Luo, Alan L. YuilleICCV 2019 · 170 citations
- Improving Video Model Transfer with Dynamic Representation LearningYi Li, Nuno VasconcelosCVPR 2022 · 2 citations
- ACTION-Net: Multipath Excitation for Action RecognitionZhengwei Wang, Qi She, Aljosa SmolicCVPR 2021
