Doodle to Detect: A Goofy but Powerful Approach to Skeleton-based Hand Gesture Recognition
Sang Hoon Han, Seonho Lee, Hyeok Nam, Jae Hyeon Park, Minhee Cha, Min Geol Kim, Hyunse Lee, Sangyeon Ahn, Moon Ju Chae, Sung In Cho
Abstract
Skeleton-based hand gesture recognition plays a crucial role in enabling intuitive human-computer interaction. Traditional methods have primarily relied on hand-crafted features-such as distances between joints or positional changes across frames-to alleviate issues from viewpoint variation or body proportion differences. However, these hand-crafted features often fail to capture the full spatio-temporal information in raw skeleton data, exhibit poor interpretability, and depend heavily on dataset-specific preprocessing, limiting generalization. In addition, normalization strategies in traditional methods, which rely on training data, can introduce domain gaps between training and testing environments, further hindering robustness in diverse real-world settings. To overcome these challenges, we exclude traditional hand-crafted features and propose Skeleton Kinematics Extraction Through Coordinated grapH (SKETCH), a novel framework that directly utilizes raw four-dimensional (time, x, y, and z) skeleton sequences and transforms them into intuitive visual graph representations. The proposed framework incorporates a novel learnable Dynamic Range Embedding (DRE) to preserve axis-wise motion magnitudes lost during normalization and visual graph representations, enabling richer and more discriminative feature learning. This approach produces a graph image that richly captures the raw data's inherent information and provides interpretable visual attention cues. Furthermore, SKETCH applies independent min-max normalization on fixed-length temporal windows in real time, mitigating degradation from absolute coordinate fluctuations caused by varying sensor viewpoints or differences in individual body proportions. Through these designs, our approach becomes inherently topology-agnostic, avoiding fragile dependencies on dataset-or sensor-specific skeleton definitions. By leveraging pre-trained vision backbones, SKETCH achieves efficient convergence and superior recognition accuracy. Experimental results on SHREC'19 and SHREC'22 benchmarks show that it outperforms state-of-the-art methods in both robustness and generalization, establishing a new paradigm for skeleton-based hand gesture recognition. The code is available at https://github.com/capableofanything/SKETCH.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2075cd09-a22e-4de0-973b-63beb36b289aBuilds on19
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Point Transformer V2: Grouped Vector Attention and Partition-based PoolingXiaoyang Wu, Yixing Lao, Li Jiang, Xihui Liu et al.NeurIPS 2022 · 924 citations
- Point-BERT: Pre-training 3D Point Cloud Transformers with Masked Point ModelingXumin Yu, Lulu Tang, Yongming Rao, Tiejun Huang et al.CVPR 2022 · 684 citations
- InfoGCN: Representation Learning for Human Skeleton-based Action RecognitionHyung-Gun Chi, Myoung Hoon Ha, Seung-geun Chi, Sang Wan Lee et al.CVPR 2022 · 383 citations
Related papers
- SkeleTR: Towards Skeleton-based Action Recognition in the WildHaodong Duan, Mingze Xu, Bing Shuai, Davide Modolo et al.ICCV 2023 · 38 citations
- Decoupled Representation Learning for Skeleton-Based Gesture RecognitionJianbo Liu, Yongcheng Liu, Ying Wang, Véronique Prinet et al.CVPR 2020
- Skeleton Compression and Complementary Enhanced Fusion Under Branch-Stage Supervision for Human Action RecognitionQin Li, Congcong Xiao, Limei Liu, Han Peng et al.ACM MM 2025
- SkeletonMAE: Graph-based Masked Autoencoder for Skeleton Sequence Pre-trainingHong Yan, Yang Liu, Yushen Wei, Zhen Li et al.ICCV 2023 · 77 citations
- STST: Spatial-Temporal Specialized Transformer for Skeleton-based Action RecognitionYuhan Zhang, Bo Wu, Wen Li, Lixin Duan et al.ACM MM 2021 · 135 citations
