Taylor Videos for Action Recognition
Lei Wang, Xiuyuan Yuan, Tom Gedeon, Liang Zheng
Abstract
Effectively extracting motions from video is a critical and long-standing problem for action recognition. This problem is very challenging because motions (i) do not have an explicit form, (ii) have various concepts such as displacement, velocity, and acceleration, and (iii) often contain noise caused by unstable pixels. Addressing these challenges, we propose the Taylor video, a new video format that highlights the dominant motions (e.g., a waving hand) in each of its frames named the Taylor frame. Taylor video is named after Taylor series, which approximates a function at a given point using important terms. In the scenario of videos, we define an implicit motionextraction function which aims to extract motions from video temporal blocks. In these blocks, using the frames, the difference frames, and higherorder difference frames, we perform Taylor expansion to approximate this function at the starting frame. We show the summation of the higherorder terms in the Taylor series gives us dominant motion patterns, where static objects, small and unstable motions are removed. Experimentally, we show that Taylor videos are effective inputs to popular architectures including 2D CNNs, 3D CNNs, and transformers. When used individually, Taylor videos yield competitive action recognition accuracy compared to RGB videos and optical flow. When fused with RGB or optical flow videos, further accuracy improvement is achieved. Additionally, we apply Taylor video computation to human skeleton sequences, resulting in Taylor skeleton sequences that outperform the use of original skeletons for skeletonbased action recognition. Code is available at: https://github.com/LeiWangR/video-ar .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c796ccb3-2277-44c1-813a-3e63b227b8d7Cited by top-tier papers1
Ask how each one uses itBuilds on11
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 2,927 citations
- Video Swin TransformerZe Liu, Jia Ning, Yue Cao, Yixuan Wei et al.CVPR 2022 · 1,847 citations
Related papers
- Hierarchical Temporal Transformer for 3D Hand Pose Estimation and Action Recognition from Egocentric RGB VideosYilin Wen, Hao Pan, Lei Yang, Jia Pan et al.CVPR 2023
- HumMUSS: Human Motion Understanding Using State Space ModelsArnab Kumar Mondal, Stefano Alletto, Denis TomèCVPR 2024 · 6 citations
- Deformable Video TransformerJue Wang, Lorenzo TorresaniCVPR 2022 · 40 citations
- Video Modeling With Correlation NetworksHeng Wang, Du Tran, Lorenzo Torresani, Matt FeiszliCVPR 2020
- VCT: A Video Compression TransformerFabian Mentzer, George Toderici, David Minnen, Sergi Caelles et al.NeurIPS 2022 · 155 citations
