SAFARI: Speech-Associated Facial Authentication for AR/VR Settings via Robust VIbration Signatures
Tianfang Zhang, Qiufan Ji, Zhengkun Ye, Md Mojibur Rahman Redoy Akanda, Ahmed Tanvir Mahdad, Cong Shi, Yan Wang, Nitesh Saxena, Yingying Chen
摘要
In AR/VR devices, the voice interface, serving as one of the primary AR/VR control mechanisms, enables users to interact naturally using speeches (voice commands) for accessing data, controlling applications, and engaging in remote communication/meetings. Voice authentication can be adopted to protect against unauthorized speech inputs. However, existing voice authentication mechanisms are usually susceptible to voice spoofing attacks and are unreliable under the variations of phonetic content. In this work, we propose SAFARI, a spoofing-resistant and text-independent speech authentication system that can be seamlessly integrated into AR/VR voice interfaces. The key idea is to elicit phonetic-invariant biometrics from the facial muscle vibrations upon the headset. During speech production, a user's facial muscles are deformed for articulating phoneme sounds. The facial deformations associated with the phonemes are referred to as visemes. They carry rich biometrics of the wearer's muscles, tissue, and bones, which can propagate through the head and vibrate the headset. SAFARI aims to derive reliable facial biometrics from the viseme-associated facial vibrations captured by the AR/VR motion sensors. Particularly, it identifies the vibration data segments that contain rich viseme patterns (prominent visemes) less susceptible to phonetic variations. Based on the prominent visemes, SAFARI learns on the correlations among facial vibrations of different frequencies to extract biometric representations invariant to the phonetic context. The key advantages of SAFARI are that it is suitable for commodity AR/VR headsets (no additional sensors) and is resistant to voice spoofing attacks as the conductive property of the facial vibrations prevents biometric disclosure via the air media or the audio channel. To mitigate the impacts of body motions in AR/VR scenarios, we also design a generative diffusion model trained to reconstruct the viseme patterns from the data distorted by motion artifacts. We conduct extensive experiments with two representative AR/VR headsets and 35 users under various usage and attack settings. We demonstrate that SAFARI can achieve over 96% true positive rate on verifying legitimate users while successfully rejecting different kinds of spoofing attacks with over 97% true negative rates.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- AuthGlass: Benchmarking Voice Liveness Detection and Authentication on Smart Glasses via Comprehensive Acoustic FeaturesWeiye Xu, Zhang Jiang, Siqi Zheng, Xiyuxing Zhang 等UbiComp 2026
- When VR Meets BCI: (Un)Observable Brainwave-Aware Privacy Reconstruction in the Metaverse via Unrestricted Inbuilt Motion SensorsTao Ni, Zehua Sun, Qingchuan Zhao, Wei-Bin Lee 等S&P 2026
它引用的顶会 Paper11
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Hidden Voice CommandsNicholas Carlini, Pratyush Mishra, Tavish Vaidya, Yuankai Zhang 等USENIX Security 2016 · 被引用 672 次
- Hearing Your Voice is Not Enough: An Articulatory Gesture Based Liveness Detection for Voice AuthenticationLinghan Zhang, Sheng Tan, Jie YangCCS 2017 · 被引用 212 次
- VoiceLive: A Phoneme Localization based Liveness Detection for Voice Authentication on SmartphonesLinghan Zhang, Sheng Tan, Jie Yang, Yingying ChenCCS 2016 · 被引用 187 次
- AR-Diffusion: Auto-Regressive Diffusion Model for Text GenerationTong Wu, Zhihao Fan, Xiao Liu, Hai-Tao Zheng 等NeurIPS 2023 · 被引用 170 次
相关 Paper
- Face-Mic: inferring live speech and speaker identity via subtle facial dynamics captured by AR/VR motion sensorsCong Shi, Xiangyu Xu, Tianfang Zhang, Payton Walker 等MobiCom 2021 · 被引用 89 次
- Harnessing Vital Sign Vibration Harmonics for Effortless and Inbuilt XR User AuthenticationTianfang Zhang, Qiufan Ji, Md Mojibur Rahman Redoy Akanda, Zhengkun Ye 等CCS 2025
- Virtual U: Defeating Face Liveness Detection by Building Virtual Models from Your Public PhotosYi Xu, True Price, Jan-Michael Frahm, Fabian MonroseUSENIX Security 2016 · 被引用 94 次
- Low-effort VR Headset User Authentication Using Head-reverberated Sounds with Replay ResistanceRuxin Wang, Long Huang, Chen WangS&P 2023
- HT-Auth: Secure VR Headset Authentication via Subtle Head TremorsZhixiang He, Fengyuan Ran, Jing Chen, Yangyang Gu 等UbiComp 2025
