Real-Time Multimodal Fingertip Contact Detection via Depth and Motion Fusion for Vision-Based Human–Computer Interaction
Mukhiddin Toshpulatov, Wookey Lee, Suan Lee, Geehyuk Lee
Abstract
Precise fingertip contact detection is a fundamental challenge for natural and immersive virtual reality (VR) interaction, yet existing vision-based methods suffer from insufficient accuracy, with typical depth errors (12-25 mm) far exceeding the <3 mm precision required to reliably distinguish hovering from true contact. While commercial motion-capture systems provide sub-millimeter accuracy, their prohibitive cost limits accessibility. To address this gap, we present a highly accurate and cost-effective fingertip contact detection system based on a single RGB camera. We introduce a specialized dataset of 53,300 RGB-depth pairs capturing millimeter-scale hand-table typing interactions and systematically fine-tune seven stateof-the-art depth estimation architectures. This reduces the mean absolute error (MAE) by 68%, from 12.3 mm to 3.84 mm, achieving 95.9% depth accuracy (δ 1 ) and a 94.4% contact-detection F1 score. Our system enables typing speeds of 45.6 WPM with a 3.1% character error rate under controlled conditions (38.3 WPM, 5.3% CER for firstsession novice users), approaching the performance of dedicated depth-sensing hardware and commercial VR/AR input systems while requiring only a standard RGB camera. These results demonstrate that, for specialized close-range HCI tasks, domain-specific fine-tuning on targeted datasets can outperform architectural innovation, helping democratize high-precision hand tracking. The dataset, training pipeline, evaluation scripts, and pretrained models are publicly available for research use. 1 1 https : / / muxiddin19 . github . io / Multimodal - Fingertip -Contact -Detection -via -Depth -and - Motion-Fusion
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext acf05b9a-7d70-4462-84b9-e7993d3b801dBuilds on24
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao et al.NeurIPS 2024 · 2,305 citations
- UniDepth: Universal Monocular Metric Depth EstimationLuigi Piccinelli, Yung-Hsu Yang, Christos Sakaridis, Mattia Segù et al.CVPR 2024 · 122 citations
- ECoDepth: Effective Conditioning of Diffusion Models for Monocular Depth EstimationSuraj Patni, Aradhye Agarwal, Chetan AroraCVPR 2024 · 38 citations
- EgoTouch: On-Body Touch Input Using AR/VR Headset CamerasVimal Mollyn, Chris HarrisonUIST 2024 · 29 citations
Related papers
- HaloTouch: Using IR Multi-Path Interference to Support Touch Interactions with General SurfacesZiyi Xia, Xincheng Huang, Sidney S. Fels, Robert XiaoCHI 2025 · 10 citations
- CAFI-AR: Contact-aware Freehand Interaction with AR ObjectsXiao Tang, Ruihui Li, Chi-Wing FuUbiComp 2023 · 8 citations
- Palmpad: Enabling Real-Time Index-to-Palm Touch Interaction with a Single RGB CameraZhe He, Xiangyang Wang, Yuanchun Shi, Chi Hsia et al.CHI 2025 · 6 citations
- MEgATrack: monochrome egocentric articulated hand-tracking for virtual realityShangchen Han, Beibei Liu, Randi Cabezas, Christopher D. Twigg et al.SIGGRAPH 2020 · 207 citations
- TapID: Rapid Touch Interaction in Virtual Reality using Wearable SensingManuel Meier, Paul Streli, Andreas Fender, Christian HolzIEEE VR 2021 · 71 citations
