SmallBigNet: Integrating Core and Contextual Views for Video Classification
Xianhang Li, Yali Wang, Zhipeng Zhou, Yu Qiao
Abstract
Temporal convolution has been widely used for video classification. However, it is performed on spatio-temporal contexts in a limited view, which often weakens its capacity of learning video representation. To alleviate this problem, we propose a concise and novel SmallBig network, with the cooperation of small and big views. For the current time step, the small view branch is used to learn the core semantics, while the big view branch is used to capture the contextual semantics. Unlike traditional temporal convolution, the big view branch can provide the small view branch with the most activated video features from a broader 3D receptive field. Via aggregating such bigview contexts, the small view branch can learn more robust and discriminative spatio-temporal representations for video classification. Furthermore, we propose to share convolution in the small and big view branch, which improves model compactness as well as alleviates overfitting. As a result, our SmallBigNet achieves a comparable model size like 2D CNNs, while boosting accuracy like 3D CNNs. We conduct extensive experiments on the large-scale video benchmarks, e.g., Kinetics400, Something-Something V1 and V2. Our SmallBig network outperforms a number of recent state-of-the-art approaches, in terms of accuracy and/or efficiency. The codes and models will be available on https://github.com/xhl-video/SmallBigNet .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 430b8852-d625-4af3-9aa4-d94381c752edCited by top-tier papers22
- Omni-Scale CNNs: a simple and effective kernel size configuration for time series classificationWensi Tang, Guodong Long, Lu Liu, Tianyi Zhou et al.ICLR 2022 · 163 citations
- MVFNet: Multi-View Fusion Network for Efficient Video RecognitionWenhao Wu, Dongliang He, Tianwei Lin, Fu Li et al.AAAI 2021 · 87 citations
- TAda! Temporally-Adaptive Convolutions for Video UnderstandingZiyuan Huang, Shiwei Zhang, Liang Pan, Zhiwu Qing et al.ICLR 2022 · 72 citations
- CT-Net: Channel Tensorization Network for Video ClassificationKunchang Li, Xianhang Li, Yali Wang, Jun Wang et al.ICLR 2021 · 69 citations
- Stand-Alone Inter-Frame Attention in Video ModelsFuchen Long, Zhaofan Qiu, Yingwei Pan, Ting Yao et al.CVPR 2022 · 68 citations
Builds on4
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 2,049 citations
- TEINet: Towards an Efficient Architecture for Video RecognitionZhaoyang Liu, Donghao Luo, Yabiao Wang, Limin Wang et al.AAAI 2020 · 267 citations
- Multi-Agent Reinforcement Learning Based Frame Sampling for Effective Untrimmed Video RecognitionWenhao Wu, Dongliang He, Xiao Tan, Shifeng Chen et al.ICCV 2019 · 135 citations
Related papers
- UniFormer: Unified Transformer for Efficient Spatial-Temporal Representation LearningKunchang Li, Yali Wang, Peng Gao, Guanglu Song et al.ICLR 2022
- TDN: Temporal Difference Networks for Efficient Action RecognitionLimin Wang, Zhan Tong, Bin Ji, Gangshan WuCVPR 2021
- EAC-Net: Efficient and Accurate Convolutional Network for Video RecognitionBowei Jin, Zhuo XuAAAI 2020 · 2 citations
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun et al.ICCV 2021 · 2,947 citations
- DSANet: Dynamic Segment Aggregation Network for Video-Level Representation LearningWenhao Wu, Yuxiang Zhao, Yanwu Xu, Xiao Tan et al.ACM MM 2021 · 30 citations
