DeePoint: Visual Pointing Recognition and Direction Estimation
Shu Nakamura, Yasutomo Kawanishi, Shohei Nobuhara, Ko Nishino
摘要
In this paper, we realize automatic visual recognition and direction estimation of pointing. We introduce the first neural pointing understanding method based on two key contributions. The first is the introduction of a first-of-its-kind large-scale dataset for pointing recognition and direction estimation, which we refer to as the DP Dataset. DP Dataset consists of more than 2 million frames of 33 people pointing in various styles annotated for each frame with pointing timings and 3D directions. The second is DeePoint, a novel deep network model for joint recognition and 3D direction estimation of pointing. DeePoint is a Transformer-based network which fully leverages the spatio-temporal coordination of the body parts, not just the hands. Through extensive experiments, we demonstrate the accuracy and efficiency of DeePoint. We believe DP Dataset and DeePoint will serve as a sound foundation for visual human intention understanding.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Omni-MMSI: Toward Identity-attributed Social Interaction UnderstandingXinpeng Li, Bolin Lai, Hardy Chen, Shijian Deng 等CVPR 2026 · 被引用 3 次
- Plug-and-Play Clarifier: A Zero-Shot Multimodal Framework for Egocentric Intent DisambiguationSicheng Yang, Yukai Huang, Weitong Cai, Shitong Sun 等AAAI 2026
它引用的顶会 Paper4
- Gaze360: Physically Unconstrained Gaze Estimation in the WildPetr Kellnhofer, Adrià Recasens, Simon Stent, Wojciech Matusik 等ICCV 2019 · 被引用 469 次
- Recurring the Transformer for Video Action RecognitionJiewen Yang, Xingbo Dong, Liujun Liu, Chao Zhang 等CVPR 2022 · 被引用 119 次
- HandOccNet: Occlusion-Robust 3D Hand Mesh Estimation NetworkJoonKyu Park, Yeonguk Oh, Gyeongsik Moon, Hongsuk Choi 等CVPR 2022 · 被引用 116 次
- Dynamic 3D Gaze from Afar: Deep Gaze Estimation from Temporal Eye-Head-Body CoordinationSoma Nonaka, Shohei Nobuhara, Ko NishinoCVPR 2022 · 被引用 31 次
相关 Paper
- Do You See What I Am Pointing At? Gesture-Based Egocentric Video Question AnsweringYura Choi, Roy Miles, Rolandos Alexandros Potamias, Ismail Elezi 等CVPR 2026 · 被引用 1 次
- KeypointNet: A Large-Scale 3D Keypoint Dataset Aggregated From Numerous Human AnnotationsYang You, Yujing Lou, Chengkun Li, Zhoujun Cheng 等CVPR 2020
- Toward Human Deictic Gesture Target EstimationXu Cao, Pranav Virupaksha, Sangmin Lee, Bolin Lai 等NeurIPS 2025 · 被引用 3 次
- Talking With Hands 16.2M: A Large-Scale Dataset of Synchronized Body-Finger Motion and Audio for Conversational Motion Analysis and SynthesisGilwoo Lee, Zhiwei Deng, Shugao Ma, Takaaki Shiratori 等ICCV 2019 · 被引用 114 次
- Ges3ViG : Incorporating Pointing Gestures into Language-Based 3D Visual Grounding for Embodied Reference UnderstandingAtharv Mahesh Mane, Dulanga Weerakoon, Vigneshwaran Subbaraju, Sougata Sen 等CVPR 2025
