Direction-of-Voice (DoV) Estimation for Intuitive Speech Interaction with Smart Devices Ecosystems
Karan Ahuja, Andy Kong, Mayank Goel, Chris Harrison
摘要
Future homes and offices will feature increasingly dense ecosystems of IoT devices, such as smart lighting, speakers, and domestic appliances. Voice input is a natural candidate for interacting with out-of-reach and often small devices that lack full-sized physical interfaces. However, at present, voice agents generally require wake-words and device names in order to specify the target of a spoken command (e.g., 'Hey Alexa, kitchen lights to full bright-ness'). In this research, we explore whether speech alone can be used as a directional communication channel, in much the same way visual gaze specifies a focus. Instead of a device's microphones simply receiving and processing spoken commands, we suggest they also infer the Direction of Voice (DoV). Our approach innately enables voice commands with addressability (i.e., devices know if a command was directed at them) in a natural and rapid manner. We quantify the accuracy of our implementation across users, rooms, spoken phrases, and other key factors that affect performance and usability. Taken together, we believe our DoV approach demonstrates feasibility and the promise of making distributed voice interactions much more intuitive and fluid.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper6
- IMUPoser: Full-Body Pose Estimation using IMUs in Phones, Watches, and EarbudsVimal Mollyn, Riku Arakawa, Mayank Goel, Chris Harrison 等CHI 2023 · 被引用 103 次
- Pose-on-the-Go: Approximating User Pose with Smartphone Sensor Fusion and Inverse KinematicsKaran Ahuja, Sven Mayer, Mayank Goel, Chris HarrisonCHI 2021 · 被引用 40 次
- FaceOri: Tracking Head Position and Orientation Using Ultrasonic Ranging on EarphonesYuntao Wang, Jiexin Ding, Ishan Chatterjee, Farshid Salemi Parizi 等CHI 2022 · 被引用 29 次
- EmBARDiment: an Embodied AI Agent for Productivity in XRRiccardo Bovo, Steven Abreu, Karan Ahuja, Eric J. Gonzalez 等IEEE VR 2025 · 被引用 19 次
- Enabling Voice-Accompanying Hand-to-Face Gesture Recognition with Cross-Device SensingZisu Li, Chen Liang, Yuntao Wang, Yue Qin 等CHI 2023 · 被引用 18 次
相关 Paper
- Soundr: Head Position and Orientation Prediction Using a Microphone ArrayJackie (Junrui) Yang, Gaurab Banerjee, Vishesh Gupta, Monica S. Lam 等CHI 2020 · 被引用 19 次
- How Humans Naturally Refer to Targets: Understanding Multimodal Instruction Patterns in Human-Robot InteractionLesong Jia, Makayla Chang, Yu Liu, Na DuCHI 2026
- Enhancing Mobile Voice Assistants with WorldGazeSven Mayer, Gierad Laput, Chris HarrisonCHI 2020 · 被引用 65 次
- Model-based Head Orientation Estimation for Smart DevicesQiang Yang, Yuanqing ZhengUbiComp 2021 · 被引用 18 次
- Aware: Intuitive Device Activation Using Prosody for Natural Voice InteractionsXinlei Zhang, Zixiong Su, Jun RekimotoCHI 2022 · 被引用 7 次
