Robust Dual-Modal Speech Keyword Spotting for XR Headsets
Zhuojiang Cai, Yuhan Ma, Feng Lu
Abstract
While speech interaction finds widespread utility within the Extended Reality (XR) domain, conventional vocal speech keyword spotting systems continue to grapple with formidable challenges, including suboptimal performance in noisy environments, impracticality in situations requiring silence, and susceptibility to inadvertent activations when others speak nearby. These challenges, however, can potentially be surmounted through the cost-effective fusion of voice and lip movement information. Consequently, we propose a novel vocal-echoic dual-modal keyword spotting system designed for XR headsets. We devise two different modal fusion approches and conduct experiments to test the system's performance across diverse scenarios. The results show that our dual-modal system not only consistently outperforms its single-modal counterparts, demonstrating higher precision in both typical and noisy environments, but also excels in accurately identifying silent utterances. Furthermore, we have successfully applied the system in real-time demonstrations, achieving promising results. The code is available at https://github.com/caizhuojiang/VE-KWS.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 44c2b2bf-8064-48a0-9b38-97ff964a4170Builds on7
- Hands-free interaction in immersive virtual reality: A systematic reviewPedro Monteiro, Guilherme Gonçalves, Hugo Coelho, Miguel Melo et al.IEEE VR 2021 · 116 citations
- EchoSpeech: Continuous Silent Speech Recognition on Minimally-obtrusive Eyewear Powered by Acoustic SensingRuidong Zhang, Ke Li, Yihong Hao, Yufan Wang et al.CHI 2023 · 53 citations
- C-Face: Continuously Reconstructing Facial Expressions by Deep Learning Contours of the Face with Ear-mounted Miniature CamerasTuochao Chen, Benjamin Steeper, Kinan Alsheikh, Songyun Tao et al.UIST 2020 · 49 citations
- EchoWhisper: Exploring an Acoustic-based Silent Speech Interface for Smartphone UsersYang Gao, Yincheng Jin, Jiyang Li, Seokmin Choi et al.UbiComp 2020 · 44 citations
- SpeeChin: A Smart Necklace for Silent Speech RecognitionRuidong Zhang, Mingyang Chen, Benjamin Steeper, Yaxuan Li et al.UbiComp 2022 · 43 citations
Related papers
- mmMIC: Multi-modal Speech Recognition based on mmWave RadarLong Fan, Lei Xie, Xinran Lu, Yi Li et al.INFOCOM 2023 · 41 citations
- Pantœnna: Mouth pose estimation for ar/vr headsets using low-profile antenna and impedance characteristic sensingDaehwa Kim, Chris HarrisonUIST 2023 · 7 citations
- HeadsetOff: Enabling Photorealistic Video Conferencing on Economical VR HeadsetsYili Jin, Xize Duan, Fangxin Wang, Xue LiuACM MM 2024 · 5 citations
- Harnessing Vital Sign Vibration Harmonics for Effortless and Inbuilt XR User AuthenticationTianfang Zhang, Qiufan Ji, Md Mojibur Rahman Redoy Akanda, Zhengkun Ye et al.CCS 2025
- SoundLip: Enabling Word and Sentence-level Lip Interaction for Smart DevicesQian Zhang, Dong Wang, Run Zhao, Yinggang YuUbiComp 2021 · 34 citations
