From Speech-to-Spatial: Grounding Utterances on A Live Shared View with Augmented Reality
Yoonsang Kim, Divyansh Pradhan, Devshree Jadeja, Arie E. Kaufman
Abstract
We introduce Speech-to-Spatial, a referent disambiguation framework that converts verbal remote-assistance instructions into spatially grounded AR guidance. Unlike prior systems that rely on additional cues (e.g., gesture, gaze) or manual expert annotations, Speech-to-Spatial infers the intended target solely from spoken references (speech input). Motivated by our formative study of speech referencing patterns, we characterize recurring ways people specify targets (Direct Attribute, Relational, Remembrance, and Chained) and ground them to our object-centric relational graph. Given an utterance, referent cues are parsed and rendered as persistent in-situ AR visual guidance, reducing iterative micro-guidance ("a bit more to the right", "now, stop.") during remote guidance. We demonstrate the use cases of our system with remote guided assistance and intent disambiguation scenarios. Our evaluation shows that Speech-to-Spatial improves task efficiency, reduces cognitive load, and enhances usability compared to a conventional voice-only baseline, transforming disembodied verbal instruction into visually explainable, actionable guidance on a live shared view.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5124d9dd-484f-4a1f-9f64-fb4abdb4a2c2Builds on19
- A User Study on Mixed Reality Remote Collaboration with Eye Gaze and Hand Gesture SharingHuidong Bai, Prasanth Sasikumar, Jing Yang, Mark BillinghurstCHI 2020 · 210 citations
- Language Conditioned Spatial Relation Reasoning for 3D Object GroundingShizhe Chen, Pierre-Louis Guhur, Makarand Tapaswi, Cordelia Schmid et al.NeurIPS 2022 · 173 citations
- GazePointAR: A Context-Aware Multimodal Voice Assistant for Pronoun Disambiguation in Wearable Augmented RealityJaewook Lee, Jun Wang, Elizabeth Brown, Liam Chu et al.CHI 2024 · 86 citations
- Partially Blended Realities: Aligning Dissimilar Spaces for Distributed Mixed Reality MeetingsJens Emil Sloth Grønbæk, Ken Pfeuffer, Eduardo Velloso, Morten Astrup et al.CHI 2023 · 75 citations
- Design Patterns for Situated Visualization in Augmented RealityBenjamin Lee, Michael Sedlmair, Dieter SchmalstiegIEEE VIS 2023 · 72 citations
Related papers
- Do You Really Need to Know Where "That" Is? Enhancing Support for Referencing in Collaborative Mixed Reality EnvironmentsJanet G. Johnson, Danilo Gasques, Tommy Sharkey, Evan Schmitz et al.CHI 2021 · 28 citations
- SpeechLess: Micro-utterance with Personalized Spatial Memory-aware Assistant in Everyday Augmented RealityYoonsang Kim, Devshree Jadeja, Divyansh Pradhan, Yalong Yang et al.IEEE VR 2026 · 1 citation
- Using Virtual Replicas to Improve Mixed Reality Remote CollaborationHuayuan Tian, Gun A. Lee, Huidong Bai, Mark BillinghurstIEEE VR 2023 · 57 citations
- How Humans Naturally Refer to Targets: Understanding Multimodal Instruction Patterns in Human-Robot InteractionLesong Jia, Makayla Chang, Yu Liu, Na DuCHI 2026
- Evaluating Augmented Reality Landmark Cues and Frame of Reference Displays with Virtual RealityYu Zhao, Jeanine K. Stefanucci, Sarah H. Creem-Regehr, Bobby BodenheimerIEEE VR 2023 · 39 citations
