DeePoint: Visual Pointing Recognition and Direction Estimation
Shu Nakamura, Yasutomo Kawanishi, Shohei Nobuhara, Ko Nishino
Abstract
In this paper, we realize automatic visual recognition and direction estimation of pointing. We introduce the first neural pointing understanding method based on two key contributions. The first is the introduction of a first-of-its-kind large-scale dataset for pointing recognition and direction estimation, which we refer to as the DP Dataset. DP Dataset consists of more than 2 million frames of 33 people pointing in various styles annotated for each frame with pointing timings and 3D directions. The second is DeePoint, a novel deep network model for joint recognition and 3D direction estimation of pointing. DeePoint is a Transformer-based network which fully leverages the spatio-temporal coordination of the body parts, not just the hands. Through extensive experiments, we demonstrate the accuracy and efficiency of DeePoint. We believe DP Dataset and DeePoint will serve as a sound foundation for visual human intention understanding.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Omni-MMSI: Toward Identity-attributed Social Interaction UnderstandingXinpeng Li, Bolin Lai, Hardy Chen, Shijian Deng et al.CVPR 2026 · 3 citations
- Plug-and-Play Clarifier: A Zero-Shot Multimodal Framework for Egocentric Intent DisambiguationSicheng Yang, Yukai Huang, Weitong Cai, Shitong Sun et al.AAAI 2026
Builds on4
- Gaze360: Physically Unconstrained Gaze Estimation in the WildPetr Kellnhofer, Adrià Recasens, Simon Stent, Wojciech Matusik et al.ICCV 2019 · 469 citations
- Recurring the Transformer for Video Action RecognitionJiewen Yang, Xingbo Dong, Liujun Liu, Chao Zhang et al.CVPR 2022 · 119 citations
- HandOccNet: Occlusion-Robust 3D Hand Mesh Estimation NetworkJoonKyu Park, Yeonguk Oh, Gyeongsik Moon, Hongsuk Choi et al.CVPR 2022 · 116 citations
- Dynamic 3D Gaze from Afar: Deep Gaze Estimation from Temporal Eye-Head-Body CoordinationSoma Nonaka, Shohei Nobuhara, Ko NishinoCVPR 2022 · 31 citations
Related papers
- Do You See What I Am Pointing At? Gesture-Based Egocentric Video Question AnsweringYura Choi, Roy Miles, Rolandos Alexandros Potamias, Ismail Elezi et al.CVPR 2026 · 1 citation
- KeypointNet: A Large-Scale 3D Keypoint Dataset Aggregated From Numerous Human AnnotationsYang You, Yujing Lou, Chengkun Li, Zhoujun Cheng et al.CVPR 2020
- Toward Human Deictic Gesture Target EstimationXu Cao, Pranav Virupaksha, Sangmin Lee, Bolin Lai et al.NeurIPS 2025 · 3 citations
- Talking With Hands 16.2M: A Large-Scale Dataset of Synchronized Body-Finger Motion and Audio for Conversational Motion Analysis and SynthesisGilwoo Lee, Zhiwei Deng, Shugao Ma, Takaaki Shiratori et al.ICCV 2019 · 114 citations
- Ges3ViG : Incorporating Pointing Gestures into Language-Based 3D Visual Grounding for Embodied Reference UnderstandingAtharv Mahesh Mane, Dulanga Weerakoon, Vigneshwaran Subbaraju, Sougata Sen et al.CVPR 2025
