X3D: Expanding Architectures for Efficient Video Recognition
Christoph Feichtenhofer
摘要
This paper presents X3D, a family of efficient video networks that progressively expand a tiny 2D image classification architecture along multiple network axes, in space, time, width and depth. Inspired by feature selection methods in machine learning, a simple stepwise network expansion approach is employed that expands a single axis in each step, such that good accuracy to complexity trade-off is achieved. To expand X3D to a specific target complexity, we perform progressive forward expansion followed by backward contraction. X3D achieves state-of-the-art performance while requiring 4.8× and 5.5× fewer multiply-adds and parameters for similar accuracy as previous work. Our most surprising finding is that networks with high spatiotemporal resolution can perform well, while being extremely light in terms of network width and parameters. We report competitive accuracy at unprecedented efficiency on video classification and detection benchmarks. Code is available at: https: //github.com/facebookresearch/SlowFast .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper151
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun 等ICCV 2021 · 被引用 2,947 次
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 被引用 2,927 次
- Multiscale Vision TransformersHaoqi Fan, Bo Xiong, Karttikeya Mangalam, Yanghao Li 等ICCV 2021 · 被引用 1,611 次
- VATT: Transformers for Multimodal Self-Supervised Learning from Raw Video, Audio and TextHassan Akbari, Liangzhe Yuan, Rui Qian, Wei-Hong Chuang 等NeurIPS 2021 · 被引用 782 次
- Revisiting Skeleton-based Action RecognitionHaodong Duan, Yue Zhao, Kai Chen, Dahua Lin 等CVPR 2022 · 被引用 752 次
它引用的顶会 Paper8
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le 等ICCV 2019 · 被引用 9,163 次
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- Drop an Octave: Reducing Spatial Redundancy in Convolutional Neural Networks With Octave ConvolutionYunpeng Chen, Haoqi Fan, Bing Xu, Zhicheng Yan 等ICCV 2019 · 被引用 665 次
- Video Classification With Channel-Separated Convolutional NetworksDu Tran, Heng Wang, Matt Feiszli, Lorenzo TorresaniICCV 2019 · 被引用 647 次
- SCSampler: Sampling Salient Clips From Video for Efficient Action RecognitionBruno Korbar, Du Tran, Lorenzo TorresaniICCV 2019 · 被引用 257 次
相关 Paper
- CT-Net: Channel Tensorization Network for Video ClassificationKunchang Li, Xianhang Li, Yali Wang, Jun Wang 等ICLR 2021 · 被引用 69 次
- SmallBigNet: Integrating Core and Contextual Views for Video ClassificationXianhang Li, Yali Wang, Zhipeng Zhou, Yu QiaoCVPR 2020
- Rethinking Resolution in the Context of Efficient Video RecognitionChuofan Ma, Qiushan Guo, Yi Jiang, Ping Luo 等NeurIPS 2022 · 被引用 17 次
- MVFNet: Multi-View Fusion Network for Efficient Video RecognitionWenhao Wu, Dongliang He, Tianwei Lin, Fu Li 等AAAI 2021 · 被引用 87 次
- 2D or not 2D? Adaptive 3D Convolution Selection for Efficient Video RecognitionHengduo Li, Zuxuan Wu, Abhinav Shrivastava, Larry S. DavisCVPR 2021
