OMG-Bench: A New Challenging Benchmark for Skeleton-based Online Micro Hand Gesture Recognition
Haochen Chang, Pengfei Ren, Buyuan Zhang, Da Li, Tianhao Han, HaoYang ZHANG, Liang Xie, Hongbo Chen, Erwei Yin
Abstract
Online micro gesture recognition from hand skeletons is critical for VR/AR interaction but faces challenges due to limited public datasets and task-specific algorithms. Micro gestures involve subtle motion patterns, which make constructing datasets with precise skeletons and frame-level annotations difficult. To this end, we develop a multi-view self-supervised pipeline to automatically generate skeleton data, complemented by heuristic rules and expert refinement for semi-automatic annotation. Based on this pipeline, we introduce OMG-Bench, the first large-scale public benchmark for skeleton-based online micro gesture recognition. It features 40 fine-grained gesture classes with 13,948 instances across 1,272 sequences, characterized by subtle motions, rapid dynamics, and continuous execution. To tackle these challenges, we propose Hierarchical Memory-Augmented Transformer (HMATr), an end-to-end framework that unifies gesture detection and classification by leveraging hierarchical memory banks which store frame-level details and window-level semantics to preserve historical context. In addition, it employs learnable position-aware queries initialized from the memory to implicitly encode gesture positions and semantics. Experiments show that HMATr outperforms state-of-the-art methods by 7.6% in detection rate, establishing a strong baseline for online micro gesture recognition. Our code is available in Suppl. Mat. and dataset will be available later.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 964b26a1-d31f-4cd0-b976-dbae07f21e2cBuilds on16
- DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETRShilong Liu, Feng Li, Hao Zhang, Xiao Yang et al.ICLR 2022 · 1,218 citations
- Channel-wise Topology Refinement Graph Convolution for Skeleton-Based Action RecognitionYuxin Chen, Ziqi Zhang, Chunfeng Yuan, Bing Li et al.ICCV 2021 · 871 citations
- Hierarchically Decomposed Graph Convolutional Networks for Skeleton-Based Action RecognitionJungho Lee, Minhyeok Lee, Dogyoon Lee, Sangyoun LeeICCV 2023 · 236 citations
- Reconstructing Hands in 3D with TransformersGeorgios Pavlakos, Dandan Shan, Ilija Radosavovic, Angjoo Kanazawa et al.CVPR 2024 · 110 citations
- Two Heads Are Better than One: Image-Point Cloud Network for Depth-Based 3D Hand Pose EstimationPengfei Ren, Yuchen Chen, Jiachang Hao, Haifeng Sun et al.AAAI 2023 · 28 citations
Related papers
- LiveGesture: Streamable Co-Speech Gesture Generation ModelMuhammad Usama Saleem, Mayur Jagdishbhai Patel, Ekkasit Pinyoanuntapong, Zhongxing Qin et al.CVPR 2026 · 4 citations
- Kinematic Enhanced Hypergraph Convolutional Network for Skeleton-based Human Action Recognition with LLM Training GuidesNan Ma, Beining Sun, Yiheng Han, Genbao XuACM MM 2025 · 3 citations
- STMG: A Machine Learning Microgesture Recognition System for Supporting Thumb-Based VR/AR InputKenrick Kin, Chengde Wan, Ken Koh, Andrei Marin et al.CHI 2024 · 20 citations
- ARMO: Autoregressive Rigging for Multi-Category ObjectsMingze Sun, Shiwei Mao, Keyi Chen, Yurun Chen et al.ICCV 2025 · 3 citations
- SkeleTR: Towards Skeleton-based Action Recognition in the WildHaodong Duan, Mingze Xu, Bing Shuai, Davide Modolo et al.ICCV 2023 · 38 citations
