Robust Dual-Modal Speech Keyword Spotting for XR Headsets
Zhuojiang Cai, Yuhan Ma, Feng Lu
摘要
While speech interaction finds widespread utility within the Extended Reality (XR) domain, conventional vocal speech keyword spotting systems continue to grapple with formidable challenges, including suboptimal performance in noisy environments, impracticality in situations requiring silence, and susceptibility to inadvertent activations when others speak nearby. These challenges, however, can potentially be surmounted through the cost-effective fusion of voice and lip movement information. Consequently, we propose a novel vocal-echoic dual-modal keyword spotting system designed for XR headsets. We devise two different modal fusion approches and conduct experiments to test the system's performance across diverse scenarios. The results show that our dual-modal system not only consistently outperforms its single-modal counterparts, demonstrating higher precision in both typical and noisy environments, but also excels in accurately identifying silent utterances. Furthermore, we have successfully applied the system in real-time demonstrations, achieving promising results. The code is available at https://github.com/caizhuojiang/VE-KWS.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- Hands-free interaction in immersive virtual reality: A systematic reviewPedro Monteiro, Guilherme Gonçalves, Hugo Coelho, Miguel Melo 等IEEE VR 2021 · 被引用 116 次
- EchoSpeech: Continuous Silent Speech Recognition on Minimally-obtrusive Eyewear Powered by Acoustic SensingRuidong Zhang, Ke Li, Yihong Hao, Yufan Wang 等CHI 2023 · 被引用 53 次
- C-Face: Continuously Reconstructing Facial Expressions by Deep Learning Contours of the Face with Ear-mounted Miniature CamerasTuochao Chen, Benjamin Steeper, Kinan Alsheikh, Songyun Tao 等UIST 2020 · 被引用 49 次
- EchoWhisper: Exploring an Acoustic-based Silent Speech Interface for Smartphone UsersYang Gao, Yincheng Jin, Jiyang Li, Seokmin Choi 等UbiComp 2020 · 被引用 44 次
- SpeeChin: A Smart Necklace for Silent Speech RecognitionRuidong Zhang, Mingyang Chen, Benjamin Steeper, Yaxuan Li 等UbiComp 2022 · 被引用 43 次
相关 Paper
- mmMIC: Multi-modal Speech Recognition based on mmWave RadarLong Fan, Lei Xie, Xinran Lu, Yi Li 等INFOCOM 2023 · 被引用 41 次
- Pantœnna: Mouth pose estimation for ar/vr headsets using low-profile antenna and impedance characteristic sensingDaehwa Kim, Chris HarrisonUIST 2023 · 被引用 7 次
- HeadsetOff: Enabling Photorealistic Video Conferencing on Economical VR HeadsetsYili Jin, Xize Duan, Fangxin Wang, Xue LiuACM MM 2024 · 被引用 5 次
- Harnessing Vital Sign Vibration Harmonics for Effortless and Inbuilt XR User AuthenticationTianfang Zhang, Qiufan Ji, Md Mojibur Rahman Redoy Akanda, Zhengkun Ye 等CCS 2025
- SoundLip: Enabling Word and Sentence-level Lip Interaction for Smart DevicesQian Zhang, Dong Wang, Run Zhao, Yinggang YuUbiComp 2021 · 被引用 34 次
