AWSD: Adaptive Weighted Spatiotemporal Distillation for Video Representation
Mohammad Tavakolian, Hamed Rezazadegan Tavakoli, Abdenour Hadid
Abstract
We propose an Adaptive Weighted Spatiotemporal Distillation (AWSD) technique for video representation by encoding the appearance and dynamics of the videos into a single RGB image map. This is obtained by adaptively dividing the videos into small segments and comparing two consecutive segments. This allows using pre-trained models on still images for video classification while successfully capturing the spatiotemporal variations in the videos. The adaptive segment selection enables effective encoding of the essential discriminative information of untrimmed videos. Based on Gaussian Scale Mixture, we compute the weights by extracting the mutual information between two consecutive segments. Unlike pooling-based methods, our AWSD gives more importance to the frames that characterize actions or events thanks to its adaptive segment length selection. We conducted extensive experimental analysis to evaluate the effectiveness of our proposed method and compared our results against those of recent state-of-the-art methods on four benchmark datatsets, including UCF101, HMDB51, Activ-ityNet v1.3, and Maryland. The obtained results on these benchmark datatsets showed that our method significantly outperforms earlier works and sets the new state-of-the-art performance in video classification. Code is available at the project webpage: https://mohammadt68.github . io/AWSD/
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f7a1bd08-bd51-49eb-bf14-fd4be8338e7aCited by top-tier papers3
- Can an Image Classifier Suffice For Action Recognition?Quanfu Fan, Chun-Fu Chen, Rameswar PandaICLR 2022 · 39 citations
- Condensing a Sequence to One Informative Frame for Video RecognitionZhaofan Qiu, Ting Yao, Yan Shu, Chong-Wah Ngo et al.ICCV 2021 · 13 citations
- Neural Networks Are More Productive Teachers Than Human Raters: Active Mixup for Data-Efficient Knowledge Distillation From a Blackbox ModelDongdong Wang, Yandong Li, Liqiang Wang, Boqing GongCVPR 2020
Related papers
- Time-Equivariant Contrastive Video Representation LearningSimon Jenni, Hailin JinICCV 2021 · 64 citations
- Multi-Group Multi-Attention: Towards Discriminative Spatiotemporal RepresentationZhensheng Shi, Liangjie Cao, Cheng Guan, Ju Liang et al.ACM MM 2020 · 1 citation
- DistInit: Learning Video Representations Without a Single Labeled VideoRohit Girdhar, Du Tran, Lorenzo Torresani, Deva RamananICCV 2019 · 59 citations
- Probabilistic Representations for Video Contrastive LearningJungin Park, Jiyoung Lee, Ig-Jae Kim, Kwanghoon SohnCVPR 2022 · 42 citations
- Class-Incremental Learning for Action Recognition in VideosJaeyoo Park, Minsoo Kang, Bohyung HanICCV 2021 · 68 citations
