Regional Attention with Architecture-Rebuilt 3D Network for RGB-D Gesture Recognition
Benjia Zhou, Yunan Li, Jun Wan
Abstract
Human gesture recognition has drawn much attention in the area of computer vision. However, the performance of gesture recognition is always influenced by some gesture-irrelevant factors like the background and the clothes of performers. Therefore, focusing on the regions of hand/arm is important to the gesture recognition. Meanwhile, a more adaptive architecture-searched network structure can also perform better than the block-fixed ones like ResNet since it increases the diversity of features in different stages of the network better. In this paper, we propose a regional attention with architecture-rebuilt 3D network (RAAR3DNet) for gesture recognition. We replace the fixed Inception modules with the automatically rebuilt structure through the network via Neural Architecture Search (NAS), owing to the different shape and representation ability of features in the early, middle, and late stage of the network. It enables the network to capture different levels of feature representations at different layers more adaptively. Meanwhile, we also design a stackable regional attention module called Dynamic-Static Attention (DSA), which derives a Gaussian guidance heatmap and dynamic motion map to highlight the hand/arm regions and the motion information in the spatial and temporal domains, respectively. Extensive experiments on two recent large-scale RGB-D gesture datasets validate the effectiveness of the proposed method and show it outperforms state-of-the-art methods. The codes of our method are available at: https://github.com/zhoubenjia/RAAR3DNet .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 16e4bbd1-db29-4f65-9fa0-0c156c48adc4Cited by top-tier papers5
- Gloss-free Sign Language Translation: Improving from Visual-Language PretrainingBenjia Zhou, Zhigang Chen, Albert Clapés, Jun Wan et al.ICCV 2023 · 123 citations
- Nested Collaborative Learning for Long-Tailed Visual RecognitionJun Li, Zichang Tan, Jun Wan, Zhen Lei et al.CVPR 2022 · 95 citations
- Decoupling and Recoupling Spatiotemporal Representation for RGB-D-based Motion RecognitionBenjia Zhou, Pichao Wang, Jun Wan, Yanyan Liang et al.CVPR 2022 · 43 citations
- Multi-stage Factorized Spatio-Temporal Representation for RGB-D Action and Gesture RecognitionYujun Ma, Benjia Zhou, Ruili Wang, Pichao WangACM MM 2023 · 21 citations
- Learning Robust Representations with Information Bottleneck and Memory Network for RGB-D-based Gesture RecognitionYunan Li, Huizhou Chen, Guanwen Feng, Qiguang MiaoICCV 2023 · 10 citations
Builds on2
Related papers
- Automatic Network Architecture Search for RGB-D Semantic SegmentationWenna Wang, Tao Zhuo, Xiuwei Zhang, Mingjun Sun et al.ACM MM 2023 · 6 citations
- Depth-Induced Multi-Scale Recurrent Attention Network for Saliency DetectionYongri Piao, Wei Ji, Jingjing Li, Miao Zhang et al.ICCV 2019 · 450 citations
- Dynamic Heterogeneous Graph Attention Neural Architecture SearchZeyang Zhang, Ziwei Zhang, Xin Wang, Yijian Qin et al.AAAI 2023 · 44 citations
- Dynamic Multiscale Graph Neural Networks for 3D Skeleton Based Human Motion PredictionMaosen Li, Siheng Chen, Yangheng Zhao, Ya Zhang et al.CVPR 2020
- No Pain, Big Gain: Classify Dynamic Point Cloud Sequences with Static Models by Fitting Feature-level Space-time SurfacesJia-Xing Zhong, Kaichen Zhou, Qingyong Hu, Bing Wang et al.CVPR 2022 · 22 citations
