SocialGesture: Delving into Multi-person Gesture Understanding
Xu Cao, Pranav Virupaksha, Wenqi Jia, Bolin Lai, Fiona Ryan, Sangmin Lee, James M. Rehg
Abstract
Previous research in human gesture recognition has largely overlooked multi-person interactions, which are crucial for understanding the social context of naturally occurring gestures. This limitation in existing datasets presents a significant challenge in aligning human gestures with other modalities like language and speech. To address this issue, we introduce SocialGesture, the first largescale dataset specifically designed for multi-person gesture analysis. SocialGesture features a diverse range of natural scenarios and supports multiple gesture analysis tasks, including video-based recognition and temporal localization, providing a valuable resource for advancing the study of gesture during complex social interactions. Furthermore, we propose a novel visual question answering (VQA) task to benchmark vision language models' (VLMs) performance on social gesture understanding. Our findings highlight several limitations of current gesture recognition models, offering insights into future directions for improvement in this field. SocialGesture is available at huggingface.co/datasets/IrohXu/SocialGesture.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 756a15e5-3471-44be-80ce-a318d3a8207eCited by top-tier papers3
- Multi-speaker Attention Alignment for Multimodal Social InteractionLiangyang Ouyang, Yifei Huang, Mingfang Zhang, Caixin Kang et al.CVPR 2026 · 8 citations
- Omni-MMSI: Toward Identity-attributed Social Interaction UnderstandingXinpeng Li, Bolin Lai, Hardy Chen, Shijian Deng et al.CVPR 2026 · 3 citations
- Toward Human Deictic Gesture Target EstimationXu Cao, Pranav Virupaksha, Sangmin Lee, Bolin Lai et al.NeurIPS 2025 · 3 citations
Builds on21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
Related papers
- EQA-MX: Embodied Question Answering using Multimodal ExpressionMd Mofijul Islam, Alexi Gladstone, Riashat Islam, Tariq IqbalICLR 2024 · 18 citations
- Gestura: A LVLM-Powered System Bridging Motion and Semantics for Real-Time Free-Form Gesture UnderstandingZhuoming Li, Aitong Liu, Mengxi Jia, Yubo Lu et al.UbiComp 2026 · 1 citation
- MTGS: A Novel Framework for Multi-Person Temporal Gaze Following and Social Gaze PredictionAnshul Gupta, Samy Tafasca, Arya Farkhondeh, Pierre Vuillecard et al.NeurIPS 2024 · 24 citations
- Understanding Co-Speech Gestures in-the-WildSindhu B. Hegde, K. R. Prajwal, Taein Kwon, Andrew ZissermanICCV 2025 · 4 citations
- Do You See What I Am Pointing At? Gesture-Based Egocentric Video Question AnsweringYura Choi, Roy Miles, Rolandos Alexandros Potamias, Ismail Elezi et al.CVPR 2026 · 1 citation
