Condensing a Sequence to One Informative Frame for Video Recognition
Zhaofan Qiu, Ting Yao, Yan Shu, Chong-Wah Ngo, Tao Mei
Abstract
Video is complex due to large variations in motion and rich content in fine-grained visual details. Abstracting useful information from such information-intensive media requires exhaustive computing resources. This paper studies a two-step alternative that first condenses the video sequence to an informative "frame" and then exploits off-the-shelf image recognition system on the synthetic frame. A valid question is how to define "useful information" and then distill it from a video sequence down to one synthetic frame. This paper presents a novel Informative Frame Synthesis (IFS) architecture that incorporates three objective tasks, i.e., appearance reconstruction, video categorization, motion estimation, and two regularizers, i.e., adversarial learning, color consistency. Each task equips the synthetic frame with one ability, while each regularizer enhances its visual quality. With these, by jointly learning the frame synthesis in an end-to-end manner, the generated frame is expected to encapsulate the required spatio-temporal information useful for video analysis. Extensive experiments are conducted on the large-scale Kinetics dataset. When comparing to baseline methods that map video sequence to a single image, IFS shows superior performance. More remarkably, IFS consistently demonstrates evident improvements on image-based 2D networks and clip-based 3D networks, and achieves comparable performance with the state-of-the-art methods with less computational cost.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e6e99328-80f6-4668-8447-0dfe28bb2bdfCited by top-tier papers3
- Can an Image Classifier Suffice For Action Recognition?Quanfu Fan, Chun-Fu Chen, Rameswar PandaICLR 2022 · 39 citations
- Learning a Condensed Frame for Memory-Efficient Video Class-Incremental LearningYixuan Pei, Zhiwu Qing, Jun Cen, Xiang Wang et al.NeurIPS 2022 · 22 citations
- MLP-3D: A MLP-like 3D Architecture with Grouped Time MixingZhaofan Qiu, Ting Yao, Chong-Wah Ngo, Tao MeiCVPR 2022 · 18 citations
Builds on6
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- Video Classification With Channel-Separated Convolutional NetworksDu Tran, Heng Wang, Matt Feiszli, Lorenzo TorresaniICCV 2019 · 647 citations
- STM: SpatioTemporal and Motion Encoding for Action RecognitionBoyuan Jiang, Mengmeng Wang, Weihao Gan, Wei Wu et al.ICCV 2019 · 442 citations
- Action Recognition With Spatial-Temporal Discriminative Filter BanksBrais Martínez, Davide Modolo, Yuanjun Xiong, Joseph TigheICCV 2019 · 70 citations
- AWSD: Adaptive Weighted Spatiotemporal Distillation for Video RepresentationMohammad Tavakolian, Hamed Rezazadegan Tavakoli, Abdenour HadidICCV 2019 · 7 citations
Related papers
- OCSampler: Compressing Videos to One Clip with Single-step SamplingJintao Lin, Haodong Duan, Kai Chen, Dahua Lin et al.CVPR 2022 · 27 citations
- Video Modeling With Correlation NetworksHeng Wang, Du Tran, Lorenzo Torresani, Matt FeiszliCVPR 2020
- Knowledge Integration Networks for Action RecognitionShiwen Zhang, Sheng Guo, Limin Wang, Weilin Huang et al.AAAI 2020 · 20 citations
- 2D or not 2D? Adaptive 3D Convolution Selection for Efficient Video RecognitionHengduo Li, Zuxuan Wu, Abhinav Shrivastava, Larry S. DavisCVPR 2021
- DynamoNet: Dynamic Action and Motion NetworkAli Diba, Vivek Sharma, Luc Van Gool, Rainer StiefelhagenICCV 2019 · 123 citations
