StegoType: Surface Typing from Egocentric Cameras
Mark Richardson, Fadi Botros, Yangyang Shi, Pinhao Guo, Bradford J. Snow, Linguang Zhang, Jingming Dong, Keith Vertanen, Shugao Ma, Robert Wang
Abstract
Text input is a critical component of any general purpose computing system, yet efficient and natural text input remains a challenge in AR and VR. Headset based hand-tracking has recently become pervasive among consumer VR devices and affords the opportunity to enable touch typing on virtual keyboards. We present an approach for decoding touch typing on uninstrumented flat surfaces using only egocentric camera-based hand-tracking as input. While egocentric hand-tracking accuracy is limited by issues like self occlusion and image fidelity, we show that a sufficiently diverse training set of hand motions paired with typed text can enable a deep learning model to extract signal from this noisy input. Furthermore, by carefully designing a closed-loop data collection process, we can train an end-to-end text decoder that accounts for natural sloppy typing on virtual keyboards. We evaluate our work with a user study (n=18) showing a mean online throughput of 42.4 WPM with an uncorrected error rate (UER) of 7% with our method compared to a physical keyboard baseline of 74.5 WPM at 0.8% UER, showing progress towards unlocking productivity and high throughput use cases in AR/VR.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d5c7114c-db0c-4ca8-9e7e-c952d4750122Cited by top-tier papers5
- TouchInsight: Uncertainty-aware Rapid Touch and Text Input for Mixed Reality from Egocentric VisionPaul Streli, Mark Richardson, Fadi Botros, Shugao Ma et al.UIST 2024 · 22 citations
- FineType: Fine-grained Tapping Gesture Recognition for Text EntryChentao Li, Ziheng Xi, Jianjiang Feng, Jie ZhouCHI 2025 · 5 citations
- EgoPressure: A Dataset for Hand Pressure and Pose Estimation in Egocentric VisionYiming Zhao, Taein Kwon, Paul Streli, Marc Pollefeys et al.CVPR 2025
- Real-Time Multimodal Fingertip Contact Detection via Depth and Motion Fusion for Vision-Based Human–Computer InteractionMukhiddin Toshpulatov, Wookey Lee, Suan Lee, Geehyuk LeeCVPR 2026
- SurfaceXR: Fusing Smartwatch IMUs and Egocentric Hand Pose for Seamless Surface InteractionsVasco Xu, Brian Chen, Eric J. Gonzalez, Andrea Colaço et al.IEEE VR 2026
Builds on10
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun et al.ICCV 2021 · 2,947 citations
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 2,927 citations
- Revisiting Skeleton-based Action RecognitionHaodong Duan, Yue Zhao, Kai Chen, Dahua Lin et al.CVPR 2022 · 752 citations
- Assembly101: A Large-Scale Multi-View Video Dataset for Understanding Procedural ActivitiesFadime Sener, Dibyadip Chatterjee, Daniel Shelepov, Kun He et al.CVPR 2022 · 168 citations
Related papers
- Decoding Surface Touch Typing from Hand-TrackingMark Richardson, Matt Durasoff, Robert WangUIST 2020 · 40 citations
- TypeAnywhere: A QWERTY-Based Text Entry Solution for Ubiquitous ComputingMingrui Ray Zhang, Shumin Zhai, Jacob O. WobbrockCHI 2022 · 27 citations
- Text Entry in Virtual Environments using Speech and a Midair KeyboardJiban Adhikary, Keith VertanenIEEE VR 2021 · 59 citations
- SkiMR: Dwell-free Eye Typing in Mixed RealityJinghui Hu, John J. Dudley, Per Ola KristenssonIEEE VR 2024 · 6 citations
- FanType: Intention-Inferring Fan-Shaped Thumb Interface for Text Entry on Small XR KeyboardsGuanghan Zhao, Louis Teys, Gyeonghwan Yang, Shengdong Zhao et al.IEEE VR 2026
