DynamoNet: Dynamic Action and Motion Network
Ali Diba, Vivek Sharma, Luc Van Gool, Rainer Stiefelhagen
Abstract
In this paper, we are interested in self-supervised learning the motion cues in videos using dynamic motion filters for a better motion representation to finally boost human action recognition in particular. Thus far, the vision community has focused on spatio-temporal approaches using standard filters, rather we here propose dynamic filters that adaptively learn the video-specific internal motion representation by predicting the short-term future frames. We name this new motion representation, as dynamic motion representation (DMR) and is embedded inside of 3D convolutional network as a new layer, which captures the visual appearance and motion dynamics throughout entire video clip via end-to-end network learning. Simultaneously, we utilize these motion representation to enrich video classification. We have designed the frame prediction task as an auxiliary task to empower the classification problem. With these overall objectives, to this end, we introduce a novel unified spatio-temporal 3D-CNN architecture (Dy-namoNet) that jointly optimizes the video classification and learning motion representation by predicting future frames as a multi-task learning problem. We conduct experiments on challenging human action datasets: Kinetics 400, UCF101, HMDB51. The experiments using the proposed DynamoNet show promising results on all the datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bfb96bc5-44c9-4a64-81f7-4689991c3e47Cited by top-tier papers31
- Masked Autoencoders As Spatiotemporal LearnersChristoph Feichtenhofer, Haoqi Fan, Yanghao Li, Kaiming HeNeurIPS 2022 · 690 citations
- Omni-Dimensional Dynamic ConvolutionChao Li, Aojun Zhou, Anbang YaoICLR 2022 · 408 citations
- Self-supervised Co-Training for Video Representation LearningTengda Han, Weidi Xie, Andrew ZissermanNeurIPS 2020 · 405 citations
- Dynamic Context-Sensitive Filtering Network for Video Salient Object DetectionMiao Zhang, Jie Liu, Yifei Wang, Yongri Piao et al.ICCV 2021 · 112 citations
- Vi2CLR: Video and Image for Visual Contrastive Learning of RepresentationAli Diba, Vivek Sharma, Reza Safdari, Dariush Lotfi et al.ICCV 2021 · 65 citations
Related papers
- Self-Supervised Video Representation Learning by Context and Motion DecouplingLianghua Huang, Yu Liu, Bin Wang, Pan Pan et al.CVPR 2021
- Video Modeling With Correlation NetworksHeng Wang, Du Tran, Lorenzo Torresani, Matt FeiszliCVPR 2020
- STM: SpatioTemporal and Motion Encoding for Action RecognitionBoyuan Jiang, Mengmeng Wang, Weihao Gan, Wei Wu et al.ICCV 2019 · 442 citations
- Self-Supervised Video Representation Learning via Latent Time NavigationDi Yang, Yaohui Wang, Quan Kong, Antitza Dantcheva et al.AAAI 2023 · 18 citations
- Motion-Augmented Self-Training for Video Recognition at Smaller ScaleKirill Gavrilyuk, Mihir Jain, Ilia Karmanov, Cees G. M. SnoekICCV 2021 · 25 citations
