MVFNet: Multi-View Fusion Network for Efficient Video Recognition
Wenhao Wu, Dongliang He, Tianwei Lin, Fu Li, Chuang Gan, Errui Ding
摘要
Conventionally, spatiotemporal modeling network and its complexity are the two most concentrated research topics in video action recognition. Existing state-of-the-art methods have achieved excellent accuracy regardless of the complexity meanwhile efficient spatiotemporal modeling solutions are slightly inferior in performance. In this paper, we attempt to acquire both efficiency and effectiveness simultaneously. First of all, besides traditionally treating H ×W ×T video frames as space-time signal (viewing from the Height-Width spatial plane), we propose to also model video from the other two Height-Time and Width-Time planes, to capture the dynamics of video thoroughly. Secondly, our model is designed based on 2D CNN backbones and model complexity is well kept in mind by design. Specifically, we introduce a novel multi-view fusion (MVF) module to exploit video dynamics using separable convolution for efficiency. It is a plug-and-play module and can be inserted into off-theshelf 2D CNNs to form a simple yet effective model called MVFNet. Moreover, MVFNet can be thought of as a generalized video modeling framework and it can specialize to be existing methods such as C2D, SlowOnly, and TSM under different settings. Extensive experiments are conducted on popular benchmarks (i.e., Something-Something V1 & V2, Kinetics, UCF-101, and HMDB-51) to show its superiority. The proposed MVFNet can achieve state-of-the-art performance but maintain 2D CNN's complexity.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Revisiting Classifier: Transferring Vision-Language Models for Video RecognitionWenhao Wu, Zhun Sun, Wanli OuyangAAAI 2023 · 被引用 141 次
- MGSampler: An Explainable Sampling Strategy for Video Action RecognitionYuan Zhi, Zhan Tong, Limin Wang, Gangshan WuICCV 2021 · 被引用 89 次
- ASCNet: Self-supervised Video Representation Learning with Appearance-Speed ConsistencyDeng Huang, Wenhao Wu, Weiwen Hu, Xu Liu 等ICCV 2021 · 被引用 55 次
- DSANet: Dynamic Segment Aggregation Network for Video-Level Representation LearningWenhao Wu, Yuxiang Zhao, Yanwu Xu, Xiao Tan 等ACM MM 2021 · 被引用 30 次
- Temporal Action Proposal Generation with Background ConstraintHaosen Yang, Wenhao Wu, Lining Wang, Sheng Jin 等AAAI 2022 · 被引用 29 次
它引用的顶会 Paper14
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 被引用 2,049 次
- Video Classification With Channel-Separated Convolutional NetworksDu Tran, Heng Wang, Matt Feiszli, Lorenzo TorresaniICCV 2019 · 被引用 647 次
- STM: SpatioTemporal and Motion Encoding for Action RecognitionBoyuan Jiang, Mengmeng Wang, Weihao Gan, Wei Wu 等ICCV 2019 · 被引用 442 次
- Self-supervised Co-Training for Video Representation LearningTengda Han, Weidi Xie, Andrew ZissermanNeurIPS 2020 · 被引用 405 次
- TEINet: Towards an Efficient Architecture for Video RecognitionZhaoyang Liu, Donghao Luo, Yabiao Wang, Limin Wang 等AAAI 2020 · 被引用 267 次
相关 Paper
- TDN: Temporal Difference Networks for Efficient Action RecognitionLimin Wang, Zhan Tong, Bin Ji, Gangshan WuCVPR 2021
- ACTION-Net: Multipath Excitation for Action RecognitionZhengwei Wang, Qi She, Aljosa SmolicCVPR 2021
- EAC-Net: Efficient and Accurate Convolutional Network for Video RecognitionBowei Jin, Zhuo XuAAAI 2020 · 被引用 2 次
- CT-Net: Channel Tensorization Network for Video ClassificationKunchang Li, Xianhang Li, Yali Wang, Jun Wang 等ICLR 2021 · 被引用 69 次
- AdaFuse: Adaptive Temporal Fusion Network for Efficient Action RecognitionYue Meng, Rameswar Panda, Chung-Ching Lin, Prasanna Sattigeri 等ICLR 2021 · 被引用 70 次
