CCS2026

Turning Everyday Earphones into a Full-Duplex Speech Eavesdropping Platform via Zero-permission IMU

Ming Gao, Ayijiaken Amantai, Yichen Dai, Jia Lv, Kaiyan Cui, Jinsong Han, Minhao Cui, Fu Xiao

摘要

With the rapid development of wearable electronics, commercial off-the-shelf earphones have been widely adopted in daily life. Inertial Measurement Units (IMUs) are standard hardware on modern earphones, which are originally designed for motion detection and user interaction. When sound propagates through the earphone structure, it will induce tiny physical vibrations, which can be recorded by built-in IMUs. Such physical responses can be exploited as a side channel to recover speech content, leading to serious privacy threats. Different from traditional audio eavesdropping methods, this work focuses on zero-permission IMU side-channel attacks, which requires no audio access or privileged permissions from the operating system. Considering the practical constraints including ultra-low sampling rate (25 Hz) and strong motion noise in real-world deployment, we propose an integrated solution combining multi-axial sensor data fusion, dual-stream attention networks, feature pyramid structures and conditional generative adversarial networks (CGAN). The system supports full-duplex and open-vocabulary speech recognition, and is capable of separating and recognizing mixed signals from human speech and on-device audio playback. This repository releases the complete implementation of the proposed framework, including raw data processing pipelines, feature extraction modules, inference engines, image denoising and cross-sensor conversion tools. We also release near-end and far-end IMU samples, category label files and well-trained model weights for validation. Due to privacy protection regulations, full datasets are not publicly available. The released codes and resources can be used for algorithm verification, performance evaluation and security research on wearable device side-channel vulnerabilities.