Exploring Acoustic Reverse Nonlinearity Against Speech Forgery in Real-Time Voice Applications
Ming Gao, Lingfeng Zhang, Yike Chen, Sifeng He, Feng Qian, Lei Yang, Fu Xiao, Jinsong Han
Abstract
Unauthorized editing of speech recordings poses a significant threat to the security and authenticity of speeches, particularly in the forensic and legal fields. Even worse, the speech is increasingly at risk of being tampered with due to the development of AI techniques (e.g., Audio Deepfake). It is difficult for normal users to guarantee what they say has not been illegally changed. Audio watermark techniques are recognized as an active method against speech forgery. However, such techniques suffer from audio quality degradation and non-real-time insertion. Therefore, they cannot be adopted into real-time voice applications against forgery on remote recordings, e.g., phone calls, live broadcasts, and online meetings. Fortunately, high-definition (HD) audio techniques provide ultrasonic bands without distortion. Therefore, ultrasonic creditable factors can be utilized. We propose an audio tamper-proof system, named Aegis. It provides commodity mobile devices (e.g., smartphones) with an effective method of real-time insertion of inaudible creditable factors. Users can claim that audio with no or mismatched ultrasound is invalid and illegal. In particular, we explore the acoustic reverse-nonlinear phenomenon where audible signals can be modulated onto the ultrasonic spectrum. By emphasizing the correlation between speech signals and ultrasound, we realize effective defense against various tampering methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 99e8df05-938c-4d7f-be25-455c138a9540Builds on9
- DolphinAttack: Inaudible Voice CommandsGuoming Zhang, Chen Yan, Xiaoyu Ji, Tianchen Zhang et al.CCS 2017 · 753 citations
- Hearing Your Voice is Not Enough: An Articulatory Gesture Based Liveness Detection for Voice AuthenticationLinghan Zhang, Sheng Tan, Jie YangCCS 2017 · 212 citations
- UltraSE: single-channel speech enhancement using ultrasoundKe Sun, Xinyu ZhangMobiCom 2021 · 69 citations
- Wearable Microphone JammingYuxin Chen, Huiying Li, Shan-Yuan Teng, Steven Nagels et al.CHI 2020 · 62 citations
- VocalLock: Sensing Vocal Tract for Passphrase-Independent User Authentication Leveraging Acoustic Signals on SmartphonesLi Lu, Jiadi Yu, Yingying Chen, Yan WangUbiComp 2020 · 35 citations
Related papers
- EchoFence: Non-Intrusive Forgery Detection in Video Conferencing via Ultrasonic SensingLeqi Zhao, Luxin Shi, Jianwei Liu, Rui Xiao et al.INFOCOM 2026
- AudioMarkNet: Audio Watermarking for Deepfake Speech DetectionWei Zong, Yang-Wai Chow, Willy Susilo, Joonsang Baek et al.USENIX Security 2025
- DeAR: A Deep-Learning-Based Audio Re-recording Resilient WatermarkingChang Liu, Jie Zhang, Han Fang, Zehua Ma et al.AAAI 2023 · 67 citations
- SafeSpeech: Robust and Universal Voice Protection Against Malicious Speech SynthesisZhisheng Zhang, Derui Wang, Qianyi Yang, Pengyang Huang et al.USENIX Security 2025
- EditGuard: Versatile Image Watermarking for Tamper Localization and Copyright ProtectionXuanyu Zhang, Runyi Li, Jiwen Yu, Youmin Xu et al.CVPR 2024 · 58 citations
