2D or not 2D? Adaptive 3D Convolution Selection for Efficient Video Recognition
Hengduo Li, Zuxuan Wu, Abhinav Shrivastava, Larry S. Davis
Abstract
3D convolutional networks are prevalent for video recognition. While achieving excellent recognition performance on standard benchmarks, they operate on a sequence of frames with 3D convolutions and thus are computationally demanding. Exploiting large variations among different videos, we introduce Ada3D, a conditional computation framework that learns instance-specific 3D usage policies to determine frames and convolution layers to be used in a 3D network. These policies are derived with a twohead lightweight selection network conditioned on each input video clip. Then, only frames and convolutions that are selected by the selection network are used in the 3D model to generate predictions. The selection network is optimized with policy gradient methods to maximize a reward that encourages making correct predictions with limited computation. We conduct experiments on three video recognition benchmarks and demonstrate that our method achieves similar accuracies to state-of-the-art 3D models while requiring 20% -50% less computation across different datasets. We also show that learned policies are transferable and Ada3D is compatible to different backbones and modern clip selection approaches. Our qualitative analysis indicates that our method allocates fewer 3D convolutions and frames for "static" inputs, yet uses more for motionintensive clips.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7e7da192-c0c3-4c61-b4a0-9f83707a2a1bCited by top-tier papers10
- Adaptive Focus for Efficient Video RecognitionYulin Wang, Zhaoxi Chen, Haojun Jiang, Shiji Song et al.ICCV 2021 · 117 citations
- OCSampler: Compressing Videos to One Clip with Single-step SamplingJintao Lin, Haodong Duan, Kai Chen, Dahua Lin et al.CVPR 2022 · 27 citations
- Multi-event Video-Text RetrievalGengyuan Zhang, Jisen Ren, Jindong Gu, Volker TrespICCV 2023 · 19 citations
- SpotEM: Efficient Video Search for Episodic MemorySanthosh Kumar Ramakrishnan, Ziad Al-Halah, Kristen GraumanICML 2023 · 15 citations
- View while Moving: Efficient Video Recognition in Long-untrimmed VideosYe Tian, Mengyu Yang, Lanshan Zhang, Zhizhen Zhang et al.ACM MM 2023 · 11 citations
Builds on13
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 2,049 citations
- Video Classification With Channel-Separated Convolutional NetworksDu Tran, Heng Wang, Matt Feiszli, Lorenzo TorresaniICCV 2019 · 647 citations
- SCSampler: Sampling Salient Clips From Video for Efficient Action RecognitionBruno Korbar, Du Tran, Lorenzo TorresaniICCV 2019 · 257 citations
- Improved Techniques for Training Adaptive Deep NetworksHao Li, Hong Zhang, Xiaojuan Qi, Ruigang Yang et al.ICCV 2019 · 152 citations
Related papers
- Dynamic Network Quantization for Efficient Video InferenceXimeng Sun, Rameswar Panda, Chun-Fu (Richard) Chen, Aude Oliva et al.ICCV 2021 · 56 citations
- Optimization Planning for 3D ConvNetsZhaofan Qiu, Ting Yao, Chong-Wah Ngo, Tao MeiICML 2021 · 9 citations
- VA-RED2: Video Adaptive Redundancy ReductionBowen Pan, Rameswar Panda, Camilo Luciano Fosco, Chung-Ching Lin et al.ICLR 2021 · 20 citations
- FrameExit: Conditional Early Exiting for Efficient Video RecognitionAmir Ghodrati, Babak Ehteshami Bejnordi, Amirhossein HabibianCVPR 2021
- AdaMML: Adaptive Multi-Modal Learning for Efficient Video RecognitionRameswar Panda, Chun-Fu (Richard) Chen, Quanfu Fan, Ximeng Sun et al.ICCV 2021 · 65 citations
