Gestura: A LVLM-Powered System Bridging Motion and Semantics for Real-Time Free-Form Gesture Understanding
Zhuoming Li, Aitong Liu, Mengxi Jia, Yubo Lu, Tengxiang Zhang, Changzhi Sun, Dell Zhang, Xuelong Li
摘要
Free-form gesture understanding is highly appealing for human-computer interaction, as it liberates users from the constraints of predefined gesture categories. However, the sole existing solution—GestureGPT—suffers from limited recognition accuracy and slow response times. In this paper, we propose Gestura, an end-to-end system for free-form gesture understanding. Gestura harnesses a pre-trained Large Vision-Language Model (LVLM) to align the highly dynamic and diverse patterns of free-form gestures with high-level semantic concepts. To better capture subtle hand movements across different styles, we introduce a Landmark Processing Module that compensate for LVLMs' lack of fine-grained domain knowledge by embedding anatomical hand priors. Further, a Chain-of-Thought (CoT) reasoning strategy enables step-by-step semantic inference, transforming shallow knowledge into deep semantic understanding and significantly enhancing the model's ability to interpret ambiguous or unconventional gestures. Together, these components allow Gestura to achieve robust and adaptable free-form gesture comprehension. Additionally, we have developed the first open-source dataset for free-form gesture intention reasoning and understanding with over 300,000 annotated QA pairs. Experimental results show that Gestura achieves the accuracy of 84.73% (closed-set) / 64.14% (open-set) in the exocentric (third-person) setting and 66.14% (closed-set) / 21.71% (open-set) in the egocentric (first-person) setting, achieving approximately 20% and 40% higher accuracy on closed-set and open-set tasks, respectively, compared to GestureGPT. Moreover, Gestura achieves over a 100× speedup in response time (1.6 seconds vs. 227 seconds) on an 8B-sized model deployed on a single NVIDIA A100 40GB GPU, and has been validated through real-device experiments with an edge-cloud collaborative setup, bringing free-form gesture understanding markedly closer to practical, real-world deployment. Both the dataset and code about the project can be accessed at https://evans-lx.github.io/Gestura/.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper18
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Towards Revealing the Mystery behind Chain of Thought: A Theoretical PerspectiveGuhao Feng, Bohang Zhang, Yuntian Gu, Haotian Ye 等NeurIPS 2023 · 被引用 470 次
- Video-LLaVA: Learning United Visual Representation by Alignment Before ProjectionBin Lin, Yang Ye, Bin Zhu, Jiaxi Cui 等EMNLP 2024 · 被引用 231 次
相关 Paper
- SocialGesture: Delving into Multi-person Gesture UnderstandingXu Cao, Pranav Virupaksha, Wenqi Jia, Bolin Lai 等CVPR 2025
- Do You See What I Am Pointing At? Gesture-Based Egocentric Video Question AnsweringYura Choi, Roy Miles, Rolandos Alexandros Potamias, Ismail Elezi 等CVPR 2026 · 被引用 1 次
- Semantic Gesticulator: Semantics-Aware Co-Speech Gesture SynthesisZeyi Zhang, Tenglong Ao, Yuyao Zhang, Qingzhe Gao 等SIGGRAPH 2024 · 被引用 39 次
- HandVQA: Diagnosing and Improving Fine-Grained Spatial Reasoning about Hands in Vision-Language ModelsKhalequzzaman Chowdhury Sayem, Mubarrat Chowdhury, Yihalem Yimolal Tiruneh, Muneeb Ahmed Khan 等CVPR 2026 · 被引用 3 次
- LLM Knows Body Language, Too: Translating Speech Voices into Human GesturesChenghao Xu, Guangtao Lyu, Jiexi Yan, Muli Yang 等ACL 2024
