Enhancing Mobile Voice Assistants with WorldGaze
Sven Mayer, Gierad Laput, Chris Harrison
摘要
Contemporary voice assistants require that objects of inter-est be specified in spoken commands. Of course, users are often looking directly at the object or place of interest ? fine-grained, contextual information that is currently unused. We present WorldGaze, a software-only method for smartphones that provides the real-world gaze location of a user that voice agents can utilize for rapid, natural, and precise interactions. We achieve this by simultaneously opening the front and rear cameras of a smartphone. The front-facing camera is used to track the head in 3D, including estimating its direction vector. As the geometry of the front and back cameras are fixed and known, we can raycast the head vector into the 3D world scene as captured by the rear-facing camera. This allows the user to intuitively define an object or region of interest using their head gaze. We started our investigations with a qualitative exploration of competing methods, before developing a functional, real-time implementation. We conclude with an evaluation that shows WorldGaze can be quick and accurate, opening new multimodal gaze+voice interactions for mobile voice agents.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- GazePointAR: A Context-Aware Multimodal Voice Assistant for Pronoun Disambiguation in Wearable Augmented RealityJaewook Lee, Jun Wang, Elizabeth Brown, Liam Chu 等CHI 2024 · 被引用 86 次
- Augmented Object Intelligence with XR-ObjectsMustafa Doga Dogan, Eric J. Gonzalez, Karan Ahuja, Ruofei Du 等UIST 2024 · 被引用 64 次
- Dwell Selection with ML-based Intent Prediction Using Only Gaze DataToshiya Isomoto, Shota Yamanaka, Buntarou ShizukiUbiComp 2022 · 被引用 33 次
- FaceOri: Tracking Head Position and Orientation Using Ultrasonic Ranging on EarphonesYuntao Wang, Jiexin Ding, Ishan Chatterjee, Farshid Salemi Parizi 等CHI 2022 · 被引用 29 次
- PrISM-Q&A: Step-Aware Voice Assistant on a Smartwatch Enabled by Multimodal Procedure Tracking and Large Language ModelsRiku Arakawa, Jill Fain Lehman, Mayank GoelUbiComp 2025 · 被引用 20 次
相关 Paper
- Soundr: Head Position and Orientation Prediction Using a Microphone ArrayJackie (Junrui) Yang, Gaurab Banerjee, Vishesh Gupta, Monica S. Lam 等CHI 2020 · 被引用 19 次
- vGaze: Implicit Saliency-Aware Calibration for Continuous Gaze Tracking on Mobile DevicesSongzhou Yang, Yuan He, Meng JinINFOCOM 2021 · 被引用 11 次
- Direction-of-Voice (DoV) Estimation for Intuitive Speech Interaction with Smart Devices EcosystemsKaran Ahuja, Andy Kong, Mayank Goel, Chris HarrisonUIST 2020 · 被引用 29 次
- Using Speech to Visualise Shared Gaze Cues in MR Remote CollaborationAllison Jing, Gun A. Lee, Mark BillinghurstIEEE VR 2022 · 被引用 21 次
- Selecting Real-World Objects via User-Perspective Phone OcclusionYue Qin, Chun Yu, Wentao Yao, Jiachen Yao 等CHI 2023 · 被引用 1 次
