SAFARI: Speech-Associated Facial Authentication for AR/VR Settings via Robust VIbration Signatures
Tianfang Zhang, Qiufan Ji, Zhengkun Ye, Md Mojibur Rahman Redoy Akanda, Ahmed Tanvir Mahdad, Cong Shi, Yan Wang, Nitesh Saxena, Yingying Chen
Abstract
In AR/VR devices, the voice interface, serving as one of the primary AR/VR control mechanisms, enables users to interact naturally using speeches (voice commands) for accessing data, controlling applications, and engaging in remote communication/meetings. Voice authentication can be adopted to protect against unauthorized speech inputs. However, existing voice authentication mechanisms are usually susceptible to voice spoofing attacks and are unreliable under the variations of phonetic content. In this work, we propose SAFARI, a spoofing-resistant and text-independent speech authentication system that can be seamlessly integrated into AR/VR voice interfaces. The key idea is to elicit phonetic-invariant biometrics from the facial muscle vibrations upon the headset. During speech production, a user's facial muscles are deformed for articulating phoneme sounds. The facial deformations associated with the phonemes are referred to as visemes. They carry rich biometrics of the wearer's muscles, tissue, and bones, which can propagate through the head and vibrate the headset. SAFARI aims to derive reliable facial biometrics from the viseme-associated facial vibrations captured by the AR/VR motion sensors. Particularly, it identifies the vibration data segments that contain rich viseme patterns (prominent visemes) less susceptible to phonetic variations. Based on the prominent visemes, SAFARI learns on the correlations among facial vibrations of different frequencies to extract biometric representations invariant to the phonetic context. The key advantages of SAFARI are that it is suitable for commodity AR/VR headsets (no additional sensors) and is resistant to voice spoofing attacks as the conductive property of the facial vibrations prevents biometric disclosure via the air media or the audio channel. To mitigate the impacts of body motions in AR/VR scenarios, we also design a generative diffusion model trained to reconstruct the viseme patterns from the data distorted by motion artifacts. We conduct extensive experiments with two representative AR/VR headsets and 35 users under various usage and attack settings. We demonstrate that SAFARI can achieve over 96% true positive rate on verifying legitimate users while successfully rejecting different kinds of spoofing attacks with over 97% true negative rates.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dabc2365-0aa2-4df2-830f-04833e9ce743Cited by top-tier papers2
- AuthGlass: Benchmarking Voice Liveness Detection and Authentication on Smart Glasses via Comprehensive Acoustic FeaturesWeiye Xu, Zhang Jiang, Siqi Zheng, Xiyuxing Zhang et al.UbiComp 2026
- When VR Meets BCI: (Un)Observable Brainwave-Aware Privacy Reconstruction in the Metaverse via Unrestricted Inbuilt Motion SensorsTao Ni, Zehua Sun, Qingchuan Zhao, Wei-Bin Lee et al.S&P 2026
Builds on11
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Hidden Voice CommandsNicholas Carlini, Pratyush Mishra, Tavish Vaidya, Yuankai Zhang et al.USENIX Security 2016 · 672 citations
- Hearing Your Voice is Not Enough: An Articulatory Gesture Based Liveness Detection for Voice AuthenticationLinghan Zhang, Sheng Tan, Jie YangCCS 2017 · 212 citations
- VoiceLive: A Phoneme Localization based Liveness Detection for Voice Authentication on SmartphonesLinghan Zhang, Sheng Tan, Jie Yang, Yingying ChenCCS 2016 · 187 citations
- AR-Diffusion: Auto-Regressive Diffusion Model for Text GenerationTong Wu, Zhihao Fan, Xiao Liu, Hai-Tao Zheng et al.NeurIPS 2023 · 170 citations
Related papers
- Face-Mic: inferring live speech and speaker identity via subtle facial dynamics captured by AR/VR motion sensorsCong Shi, Xiangyu Xu, Tianfang Zhang, Payton Walker et al.MobiCom 2021 · 89 citations
- Harnessing Vital Sign Vibration Harmonics for Effortless and Inbuilt XR User AuthenticationTianfang Zhang, Qiufan Ji, Md Mojibur Rahman Redoy Akanda, Zhengkun Ye et al.CCS 2025
- Virtual U: Defeating Face Liveness Detection by Building Virtual Models from Your Public PhotosYi Xu, True Price, Jan-Michael Frahm, Fabian MonroseUSENIX Security 2016 · 94 citations
- Low-effort VR Headset User Authentication Using Head-reverberated Sounds with Replay ResistanceRuxin Wang, Long Huang, Chen WangS&P 2023
- HT-Auth: Secure VR Headset Authentication via Subtle Head TremorsZhixiang He, Fengyuan Ran, Jing Chen, Yangyang Gu et al.UbiComp 2025
