LD-ConGR: A Large RGB-D Video Dataset for Long-Distance Continuous Gesture Recognition
Dan Liu, Libo Zhang, Yanjun Wu
Abstract
Gesture recognition plays an important role in natural human-computer interaction and sign language recognition. Existing research on gesture recognition is limited to close-range interaction such as vehicle gesture control and face-to-face communication. To apply gesture recognition to long-distance interactive scenes such as meetings and smart homes, a large RGB-D video dataset LD-ConGR is established in this paper. LD-ConGR is distinguished from existing gesture datasets by its long-distance gesture collection, fine-grained annotations, and high video qual-ity. Specifically, 1) the farthest gesture provided by the LD-ConGR is captured 4m away from the camera while existing gesture datasets collect gestures within 1m from the camera; 2) besides the gesture category, the temporal segmentation of gestures and hand location are also anno-tated in LD-ConGR; 3) videos are captured at high reso-lution (1280 x 720 for color streams and 640 x 576 for depth streams) and high frame rate (30 fps). On top of the LD-ConGR, a series of experimental and studies are conducted, and the proposed gesture region estimation and key frame sampling strategies are demonstrated to be effective in dealing with long-distance gesture recognition and the uncertainty of gesture duration. The dataset and experimen-tal results presented in this paper are expected to boost the research of long-distance gesture recognition. The dataset is available at https://github.com/Diananini/LD-ConGR-CVPR2022.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fd289aa4-16f8-488d-b9ac-909bebcea518Cited by top-tier papers4
- Multi-Level Confidence Learning for Trustworthy Multimodal ClassificationXiao Zheng, Chang Tang, Zhiguo Wan, Chengyu Hu et al.AAAI 2023 · 41 citations
- Data-Free Class-Incremental Hand Gesture RecognitionShubhra Aich, Jesús Ruiz-Santaquiteria, Zhenyu Lu, Prachi Garg et al.ICCV 2023 · 14 citations
- Toward Human Deictic Gesture Target EstimationXu Cao, Pranav Virupaksha, Sangmin Lee, Bolin Lai et al.NeurIPS 2025 · 3 citations
- SocialGesture: Delving into Multi-person Gesture UnderstandingXu Cao, Pranav Virupaksha, Wenqi Jia, Bolin Lai et al.CVPR 2025
Builds on4
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- Fingerspelling Recognition in the Wild With Iterative Visual AttentionBowen Shi, Aurora Martinez Del Rio, Jonathan Keane, Diane Brentari et al.ICCV 2019 · 76 citations
- MMTM: Multimodal Transfer Module for CNN FusionHamid Reza Vaezi Joze, Amirreza Shaban, Michael L. Iuzzolino, Kazuhito KoishidaCVPR 2020
- Temporal Pyramid Network for Action RecognitionCeyuan Yang, Yinghao Xu, Jianping Shi, Bo Dai et al.CVPR 2020
Related papers
- MGR-Dark: A Large Multimodal Video Dataset and RGB-IR Benchmark for Gesture Recognition in DarknessYuanyuan Shi, Yunan Li, Siyu Liang, Huizhou Chen et al.ACM MM 2024 · 2 citations
- MAGIC: A Dataset Capturing Mid-Air Gesture Performance for Interaction and Feature AnalysisMasoumehsadat Hosseini, Dimitar Valkov, Donald Degraen, Heiko Müller et al.UbiComp 2026
- Real-time Arm Gesture Recognition in Smart Home Scenarios via Millimeter Wave SensingHaipeng Liu, Yuheng Wang, Anfu Zhou, Hanyue He et al.UbiComp 2021 · 149 citations
- Gestura: A LVLM-Powered System Bridging Motion and Semantics for Real-Time Free-Form Gesture UnderstandingZhuoming Li, Aitong Liu, Mengxi Jia, Yubo Lu et al.UbiComp 2026 · 1 citation
- HOI4D: A 4D Egocentric Dataset for Category-Level Human-Object InteractionYunze Liu, Yun Liu, Che Jiang, Kangbo Lyu et al.CVPR 2022 · 126 citations
