From Speech-to-Spatial: Grounding Utterances on A Live Shared View with Augmented Reality
Yoonsang Kim, Divyansh Pradhan, Devshree Jadeja, Arie E. Kaufman
摘要
We introduce Speech-to-Spatial, a referent disambiguation framework that converts verbal remote-assistance instructions into spatially grounded AR guidance. Unlike prior systems that rely on additional cues (e.g., gesture, gaze) or manual expert annotations, Speech-to-Spatial infers the intended target solely from spoken references (speech input). Motivated by our formative study of speech referencing patterns, we characterize recurring ways people specify targets (Direct Attribute, Relational, Remembrance, and Chained) and ground them to our object-centric relational graph. Given an utterance, referent cues are parsed and rendered as persistent in-situ AR visual guidance, reducing iterative micro-guidance ("a bit more to the right", "now, stop.") during remote guidance. We demonstrate the use cases of our system with remote guided assistance and intent disambiguation scenarios. Our evaluation shows that Speech-to-Spatial improves task efficiency, reduces cognitive load, and enhances usability compared to a conventional voice-only baseline, transforming disembodied verbal instruction into visually explainable, actionable guidance on a live shared view.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper19
- A User Study on Mixed Reality Remote Collaboration with Eye Gaze and Hand Gesture SharingHuidong Bai, Prasanth Sasikumar, Jing Yang, Mark BillinghurstCHI 2020 · 被引用 210 次
- Language Conditioned Spatial Relation Reasoning for 3D Object GroundingShizhe Chen, Pierre-Louis Guhur, Makarand Tapaswi, Cordelia Schmid 等NeurIPS 2022 · 被引用 173 次
- GazePointAR: A Context-Aware Multimodal Voice Assistant for Pronoun Disambiguation in Wearable Augmented RealityJaewook Lee, Jun Wang, Elizabeth Brown, Liam Chu 等CHI 2024 · 被引用 86 次
- Partially Blended Realities: Aligning Dissimilar Spaces for Distributed Mixed Reality MeetingsJens Emil Sloth Grønbæk, Ken Pfeuffer, Eduardo Velloso, Morten Astrup 等CHI 2023 · 被引用 75 次
- Design Patterns for Situated Visualization in Augmented RealityBenjamin Lee, Michael Sedlmair, Dieter SchmalstiegIEEE VIS 2023 · 被引用 72 次
相关 Paper
- Do You Really Need to Know Where "That" Is? Enhancing Support for Referencing in Collaborative Mixed Reality EnvironmentsJanet G. Johnson, Danilo Gasques, Tommy Sharkey, Evan Schmitz 等CHI 2021 · 被引用 28 次
- SpeechLess: Micro-utterance with Personalized Spatial Memory-aware Assistant in Everyday Augmented RealityYoonsang Kim, Devshree Jadeja, Divyansh Pradhan, Yalong Yang 等IEEE VR 2026 · 被引用 1 次
- Using Virtual Replicas to Improve Mixed Reality Remote CollaborationHuayuan Tian, Gun A. Lee, Huidong Bai, Mark BillinghurstIEEE VR 2023 · 被引用 57 次
- How Humans Naturally Refer to Targets: Understanding Multimodal Instruction Patterns in Human-Robot InteractionLesong Jia, Makayla Chang, Yu Liu, Na DuCHI 2026
- Evaluating Augmented Reality Landmark Cues and Frame of Reference Displays with Virtual RealityYu Zhao, Jeanine K. Stefanucci, Sarah H. Creem-Regehr, Bobby BodenheimerIEEE VR 2023 · 被引用 39 次
