Cross-Modal Learning with 3D Deformable Attention for Action Recognition
Sangwon Kim, Dasom Ahn, ByoungChul Ko
摘要
An important challenge in vision-based action recognition is the embedding of spatiotemporal features with two or more heterogeneous modalities into a single feature. In this study, we propose a new 3D deformable transformer for action recognition with adaptive spatiotemporal receptive fields and a cross-modal learning scheme. The 3D deformable transformer consists of three attention modules: 3D deformability, local joint stride, and temporal stride attention. The two cross-modal tokens are input into the 3D deformable attention module to create a cross-attention token with a reflected spatiotemporal correlation. Local joint stride attention is applied to spatially combine attention and pose tokens. Temporal stride attention temporally reduces the number of input tokens in the attention module and supports temporal expression learning without the simultaneous use of all tokens. The deformable transformer iterates L-times and combines the last cross-modal token for classification. The proposed 3D deformable transformer was tested on the NTU60, NTU120, FineGYM, and PennAction datasets, and showed results better than or similar to pre-trained state-of-the-art methods even without a pre-training process. In addition, by visualizing important joints and correlations during action recognition through spatial joint and temporal stride attention, the possibility of achieving an explainable potential for action recognition is presented.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Multi-Modality Co-Learning for Efficient Skeleton-based Action RecognitionJinfu Liu, Chen Chen, Mengyuan LiuACM MM 2024 · 被引用 27 次
- Just Add π! Pose Induced Video Transformers for Understanding Activities of Daily LivingDominick Reilly, Srijan DasCVPR 2024 · 被引用 14 次
- Giving Meaning to Movements: Challenges and Opportunities in Expanding Communication by Pairing Unaided AAC with Speech Generated MessagesImran Kabir, Sharon Ann Redmon, Lynn R. Elko, Kevin Williams 等CHI 2026 · 被引用 1 次
- Neuron: Learning Context-Aware Evolving Representations for Zero-Shot Skeleton Action RecognitionYang Chen, Jingcai Guo, Song Guo, Dacheng TaoCVPR 2025
它引用的顶会 Paper23
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun 等ICCV 2021 · 被引用 2,947 次
相关 Paper
- Deformable Video TransformerJue Wang, Lorenzo TorresaniCVPR 2022 · 被引用 40 次
- Skeleton MixFormer: Multivariate Topology Representation for Skeleton-based Action RecognitionWentian Xin, Qiguang Miao, Yi Liu, Ruyi Liu 等ACM MM 2023 · 被引用 66 次
- Optimizing Human Pose Estimation Through Focused Human and Joint RegionsYingying Jiao, Zhigang Wang, Zhenguang Liu, Shaojing Fan 等AAAI 2025 · 被引用 4 次
- Frequency Guidance Matters: Skeletal Action Recognition by Frequency-Aware Mixed TransformerWenhan Wu, Ce Zheng, Zihao Yang, Chen Chen 等ACM MM 2024 · 被引用 16 次
- Skeletal Spatial-Temporal Semantics Guided Homogeneous-Heterogeneous Multimodal Network for Action RecognitionChenwei Zhang, Yuxuan Hu, Min Yang, Chengming Li 等ACM MM 2023 · 被引用 4 次
