DSANet: Dynamic Segment Aggregation Network for Video-Level Representation Learning
Wenhao Wu, Yuxiang Zhao, Yanwu Xu, Xiao Tan, Dongliang He, Zhikang Zou, Jin Ye, Yingying Li, Mingde Yao, Zichao Dong, Yifeng Shi
摘要
Long-range and short-range temporal modeling are two complementary and crucial aspects of video recognition. Most of the state-of-the-arts focus on short-range spatio-temporal modeling and then average multiple snippet-level predictions to yield the final video-level prediction. Thus, their video-level prediction does not consider spatio-temporal features of how video evolves along the temporal dimension. In this paper, we introduce a novel Dynamic Segment Aggregation (DSA) module to capture relationship among snippets. To be more specific, we attempt to generate a dynamic kernel for a convolutional operation to aggregate long-range temporal information among adjacent snippets adaptively. The DSA module is an efficient plug-and-play module and can be combined with the off-the-shelf clip-based models (i.e., TSM, I3D) to perform powerful long-range modeling with minimal overhead. The final video architecture, coined as DSANet. We conduct extensive experiments on several video recognition benchmarks (i.e., Mini-Kinetics-200, Kinetics-400, Something-Something V1 and ActivityNet) to show its superiority. Our proposed DSA module is shown to benefit various video recognition models significantly. For example, equipped with DSA modules, the top-1 accuracy of I3D ResNet-50 is improved from 74.9% to 78.2% on Kinetics-400. Codes are available at https://github.com/whwu95/DSANet.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Revisiting Classifier: Transferring Vision-Language Models for Video RecognitionWenhao Wu, Zhun Sun, Wanli OuyangAAAI 2023 · 被引用 141 次
- Delving into the Local: Dynamic Inconsistency Learning for DeepFake Video DetectionZhihao Gu, Yang Chen, Taiping Yao, Shouhong Ding 等AAAI 2022 · 被引用 117 次
- ASCNet: Self-supervised Video Representation Learning with Appearance-Speed ConsistencyDeng Huang, Wenhao Wu, Weiwen Hu, Xu Liu 等ICCV 2021 · 被引用 55 次
- Multi-Scale Adaptive Network for Single Image DenoisingYuanbiao Gou, Peng Hu, Jiancheng Lv, Joey Tianyi Zhou 等NeurIPS 2022 · 被引用 53 次
- Temporal Action Proposal Generation with Background ConstraintHaosen Yang, Wenhao Wu, Lining Wang, Sheng Jin 等AAAI 2022 · 被引用 29 次
它引用的顶会 Paper16
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 被引用 2,049 次
- CARAFE: Content-Aware ReAssembly of FEaturesJiaqi Wang, Kai Chen, Rui Xu, Ziwei Liu 等ICCV 2019 · 被引用 842 次
- Video Classification With Channel-Separated Convolutional NetworksDu Tran, Heng Wang, Matt Feiszli, Lorenzo TorresaniICCV 2019 · 被引用 647 次
- STM: SpatioTemporal and Motion Encoding for Action RecognitionBoyuan Jiang, Mengmeng Wang, Weihao Gan, Wei Wu 等ICCV 2019 · 被引用 442 次
相关 Paper
- TAM: Temporal Adaptive Module for Video RecognitionZhaoyang Liu, Limin Wang, Wayne Wu, Chen Qian 等ICCV 2021 · 被引用 356 次
- TEA: Temporal Excitation and Aggregation for Action RecognitionYan Li, Bin Ji, Xintian Shi, Jianguo Zhang 等CVPR 2020
- TDN: Temporal Difference Networks for Efficient Action RecognitionLimin Wang, Zhan Tong, Bin Ji, Gangshan WuCVPR 2021
- Shrinking Temporal Attention in Transformers for Video Action RecognitionBonan Li, Pengfei Xiong, Congying Han, Tiande GuoAAAI 2022 · 被引用 19 次
- Temporal-attentive Covariance Pooling Networks for Video RecognitionZilin Gao, Qilong Wang, Bingbing Zhang, Qinghua Hu 等NeurIPS 2021 · 被引用 33 次
