VoxPrivacy: A Benchmark for Evaluating Interactional Privacy of Speech Language Models
Yuxiang Wang, HongYu Liu, Dekun Chen, Xueyao Zhang, Zhizheng Wu
Abstract
As Speech Language Models (SLMs) transition from personal devices to shared, multi-user environments such as smart homes, a new challenge emerges: the model is expected to distinguish between users to manage information flow appropriately. Without this capability, an SLM could reveal one user’s confidential schedule to another—a privacy failure we term interactional privacy. Thus, the ability to generate speaker-aware responses becomes essential for SLM safe deployment. Current SLM benchmarks test dialogue ability but overlook speaker identity. Multi-speaker benchmarks check who said what without assessing whether SLMs adapt their responses. Privacy benchmarks focus on globally sensitive data (e.g., bank passwords) while neglecting contextually sensitive information (e.g., a user’s private appointment). To address this gap, we introduce VoxPrivacy, the first benchmark designed to evaluate interactional privacy in SLMs. VoxPrivacy spans three tiers of increasing difficulty, from following direct secrecy commands to proactively protecting privacy. Our evaluation of nine SLMs on a 32-hour bilingual dataset reveals a widespread vulnerability: most open-source models perform close to random chance (around 50% accuracy) on conditional privacy decisions, while even strong closed-source systems still fall short on proactive privacy inference. We further validate these findings on Real-VoxPrivacy, a human-recorded subset, confirming that the failures observed on synthetic data persist in real speech. We also demonstrate a viable path forward: by fine-tuning on a new 4,000-hour training set, we improve the model’s privacy-preserving capabilities while achieving fair robustness. To support future work, we are releasing the VoxPrivacy benchmark, the large-scale training set, and the fine-tuned model to help the development of safer and more context-aware SLMs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on7
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- Can LLMs Keep a Secret? Testing Privacy Implications of Language Models via Contextual Integrity TheoryNiloofar Mireshghallah, Hyunwoo Kim, Xuhui Zhou, Yulia Tsvetkov et al.ICLR 2024 · 198 citations
- Learning to Answer Questions in Dynamic Audio-Visual ScenariosGuangyao Li, Yake Wei, Yapeng Tian, Chenliang Xu et al.CVPR 2022 · 101 citations
- AVQA: A Dataset for Audio-Visual Question Answering on VideosPinci Yang, Xin Wang, Xuguang Duan, Hong Chen et al.ACM MM 2022 · 60 citations
- SpeakerLM: End-to-End Versatile Speaker Diarization and Recognition with Multimodal Large Language ModelsHan Yin, Yafeng Chen, Chong Deng, Luyao Cheng et al.AAAI 2026 · 18 citations
Related papers
- PrivLM-Bench: A Multi-level Privacy Evaluation Benchmark for Language ModelsHaoran Li, Dadi Guo, Donghao Li, Wei Fan et al.ACL 2024 · 9 citations
- MultiPriv: Benchmarking Individual-Level Privacy Reasoning in Vision-Language ModelsXiongtao Sun, HUI LI, Jiaming Zhang, Yujie Yang et al.ICML 2026 · 3 citations
- Do Vision-Language Models Respect Contextual Integrity in Location Disclosure?Ruixin Yang, Ethan Mendes, Arthur Wang, James Hays et al.ICLR 2026 · 1 citation
- Protecting Bystander Privacy via Selective Hearing in Audio LLMsXiao Zhan, Guangzhi Sun, Jose Such, Philip C. WoodlandACL 2026
- Mind the Third Eye! Benchmarking Privacy Awareness in MLLM-powered Smartphone AgentsZhixin Lin, Jungang Li, Shidong Pan, Yibo Shi et al.AAAI 2026 · 7 citations
