X3D: Expanding Architectures for Efficient Video Recognition
Christoph Feichtenhofer
Abstract
This paper presents X3D, a family of efficient video networks that progressively expand a tiny 2D image classification architecture along multiple network axes, in space, time, width and depth. Inspired by feature selection methods in machine learning, a simple stepwise network expansion approach is employed that expands a single axis in each step, such that good accuracy to complexity trade-off is achieved. To expand X3D to a specific target complexity, we perform progressive forward expansion followed by backward contraction. X3D achieves state-of-the-art performance while requiring 4.8× and 5.5× fewer multiply-adds and parameters for similar accuracy as previous work. Our most surprising finding is that networks with high spatiotemporal resolution can perform well, while being extremely light in terms of network width and parameters. We report competitive accuracy at unprecedented efficiency on video classification and detection benchmarks. Code is available at: https: //github.com/facebookresearch/SlowFast .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2054ff3f-3b42-4d6d-b7d2-e6594bccedd7Cited by top-tier papers151
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun et al.ICCV 2021 · 2,947 citations
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 2,927 citations
- Multiscale Vision TransformersHaoqi Fan, Bo Xiong, Karttikeya Mangalam, Yanghao Li et al.ICCV 2021 · 1,611 citations
- VATT: Transformers for Multimodal Self-Supervised Learning from Raw Video, Audio and TextHassan Akbari, Liangzhe Yuan, Rui Qian, Wei-Hong Chuang et al.NeurIPS 2021 · 782 citations
- Revisiting Skeleton-based Action RecognitionHaodong Duan, Yue Zhao, Kai Chen, Dahua Lin et al.CVPR 2022 · 752 citations
Builds on8
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le et al.ICCV 2019 · 9,163 citations
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- Drop an Octave: Reducing Spatial Redundancy in Convolutional Neural Networks With Octave ConvolutionYunpeng Chen, Haoqi Fan, Bing Xu, Zhicheng Yan et al.ICCV 2019 · 665 citations
- Video Classification With Channel-Separated Convolutional NetworksDu Tran, Heng Wang, Matt Feiszli, Lorenzo TorresaniICCV 2019 · 647 citations
- SCSampler: Sampling Salient Clips From Video for Efficient Action RecognitionBruno Korbar, Du Tran, Lorenzo TorresaniICCV 2019 · 257 citations
Related papers
- CT-Net: Channel Tensorization Network for Video ClassificationKunchang Li, Xianhang Li, Yali Wang, Jun Wang et al.ICLR 2021 · 69 citations
- SmallBigNet: Integrating Core and Contextual Views for Video ClassificationXianhang Li, Yali Wang, Zhipeng Zhou, Yu QiaoCVPR 2020
- Rethinking Resolution in the Context of Efficient Video RecognitionChuofan Ma, Qiushan Guo, Yi Jiang, Ping Luo et al.NeurIPS 2022 · 17 citations
- MVFNet: Multi-View Fusion Network for Efficient Video RecognitionWenhao Wu, Dongliang He, Tianwei Lin, Fu Li et al.AAAI 2021 · 87 citations
- 2D or not 2D? Adaptive 3D Convolution Selection for Efficient Video RecognitionHengduo Li, Zuxuan Wu, Abhinav Shrivastava, Larry S. DavisCVPR 2021
