Enhancing Mobile Voice Assistants with WorldGaze
Sven Mayer, Gierad Laput, Chris Harrison
Abstract
Contemporary voice assistants require that objects of inter-est be specified in spoken commands. Of course, users are often looking directly at the object or place of interest ? fine-grained, contextual information that is currently unused. We present WorldGaze, a software-only method for smartphones that provides the real-world gaze location of a user that voice agents can utilize for rapid, natural, and precise interactions. We achieve this by simultaneously opening the front and rear cameras of a smartphone. The front-facing camera is used to track the head in 3D, including estimating its direction vector. As the geometry of the front and back cameras are fixed and known, we can raycast the head vector into the 3D world scene as captured by the rear-facing camera. This allows the user to intuitively define an object or region of interest using their head gaze. We started our investigations with a qualitative exploration of competing methods, before developing a functional, real-time implementation. We conclude with an evaluation that shows WorldGaze can be quick and accurate, opening new multimodal gaze+voice interactions for mobile voice agents.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f0d9cf6a-ba41-4953-8989-3db00b0efe9aCited by top-tier papers13
- GazePointAR: A Context-Aware Multimodal Voice Assistant for Pronoun Disambiguation in Wearable Augmented RealityJaewook Lee, Jun Wang, Elizabeth Brown, Liam Chu et al.CHI 2024 · 86 citations
- Augmented Object Intelligence with XR-ObjectsMustafa Doga Dogan, Eric J. Gonzalez, Karan Ahuja, Ruofei Du et al.UIST 2024 · 64 citations
- Dwell Selection with ML-based Intent Prediction Using Only Gaze DataToshiya Isomoto, Shota Yamanaka, Buntarou ShizukiUbiComp 2022 · 33 citations
- FaceOri: Tracking Head Position and Orientation Using Ultrasonic Ranging on EarphonesYuntao Wang, Jiexin Ding, Ishan Chatterjee, Farshid Salemi Parizi et al.CHI 2022 · 29 citations
- PrISM-Q&A: Step-Aware Voice Assistant on a Smartwatch Enabled by Multimodal Procedure Tracking and Large Language ModelsRiku Arakawa, Jill Fain Lehman, Mayank GoelUbiComp 2025 · 20 citations
Related papers
- Soundr: Head Position and Orientation Prediction Using a Microphone ArrayJackie (Junrui) Yang, Gaurab Banerjee, Vishesh Gupta, Monica S. Lam et al.CHI 2020 · 19 citations
- vGaze: Implicit Saliency-Aware Calibration for Continuous Gaze Tracking on Mobile DevicesSongzhou Yang, Yuan He, Meng JinINFOCOM 2021 · 11 citations
- Direction-of-Voice (DoV) Estimation for Intuitive Speech Interaction with Smart Devices EcosystemsKaran Ahuja, Andy Kong, Mayank Goel, Chris HarrisonUIST 2020 · 29 citations
- Using Speech to Visualise Shared Gaze Cues in MR Remote CollaborationAllison Jing, Gun A. Lee, Mark BillinghurstIEEE VR 2022 · 21 citations
- Selecting Real-World Objects via User-Perspective Phone OcclusionYue Qin, Chun Yu, Wentao Yao, Jiachen Yao et al.CHI 2023 · 1 citation
