Video-to-Image Casting: A Flatting Method for Video Analysis
Xu Chen, Chenqiang Gao, Feng Yang, Xiaohan Wang, Yi Yang, Yahong Han
Abstract
Previous mainstream video analysis methods, especially 3D CNNs-based models, mainly aim to transfer frameworks from the image domain to the video domain, and they follow the regime which has been succeeded in image processing, i.e., large-scale benchmarks and deep networks. However, processing videos is still time-consuming due to the increased computational cost. In this paper, we propose to flat the video and construct a Spatio-temporal Image (STI), i.e., squeezing the temporal dimension into a spatial plane. To pursuit the video-level modeling and efficient architecture, we devise a Collective Convolution (CoConv) operation to replace the 2D convolution. With the holistic sampling strategy, this novel operation can extract the video-level spatio-temporal representation. Moreover, we ensure that each CoConv operation has the same number of parameters as the original 2D filter, thus we can utilize a 2D network equipped with CoConv to analyze videos without additional computations. To verify the effectiveness of our method for the general video analysis, we evaluate it on three typical tasks, i.e., supervised action recognition, self-supervised action recognition, and dynamic texture recognition. Extensive experimental results show that our method can achieve comparable or state-of-the-art performances on these benchmarks while using much fewer computations compared with its 3D counterpart.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 973af442-8702-4651-acdb-afab9183a2d3Cited by top-tier papers1
Ask how each one uses itRelated papers
- Grouped Spatial-Temporal Aggregation for Efficient Action RecognitionChenxu Luo, Alan L. YuilleICCV 2019 · 170 citations
- Deep Analysis of CNN-Based Spatio-Temporal Representations for Action RecognitionChun-Fu (Richard) Chen, Rameswar Panda, Kandan Ramakrishnan, Rogério Feris et al.CVPR 2021
- Spatiotemporal Joint Filter Decomposition in 3D Convolutional Neural NetworksZichen Miao, Ze Wang, Xiuyuan Cheng, Qiang QiuNeurIPS 2021 · 12 citations
- SSAN: Separable Self-Attention Network for Video Representation LearningXudong Guo, Xun Guo, Yan LuCVPR 2021
- ACTION-Net: Multipath Excitation for Action RecognitionZhengwei Wang, Qi She, Aljosa SmolicCVPR 2021
