Protecting Bystander Privacy via Selective Hearing in Audio LLMs
Xiao Zhan, Guangzhi Sun, Jose Such, Philip C. Woodland
Abstract
Audio Large language models (LLMs) are increasingly deployed in the real world, where they inevitably capture speech from unintended nearby bystanders, raising privacy risks that existing benchmarks and defences did not consider. We introduce SH-Bench, the first benchmark designed to evaluate selective hearing: a model's ability to attend to an intended main speaker while refusing to process or reveal information about incidental bystander speech. SH-Bench contains 3,968 multi-speaker audio mixtures, including both real-world and synthetic scenarios, paired with 77k multiple-choice questions that probe models under general and selective operating modes. In addition, we propose Selective Efficacy (SE), a novel metric capturing both multi-speaker comprehension and bystander-privacy protection. Our evaluation of state-of-the-art open-source and proprietary LLMs reveals substantial bystander privacy leakage, with strong audio understanding failing to translate into selective protection of bystander privacy. To mitigate this gap, we also present Bystander Privacy Fine-Tuning (BPFT), a novel training pipeline that teaches models to refuse bystander-related queries without degrading main-speaker comprehension. We show that BPFT yields substantial gains, achieving an absolute 47% higher bystander accuracy under selective mode and an absolute 16% higher SE compared to Gemini 2.5 Pro, which is the best audio LLM without BPFT. Together, SH-Bench and BPFT provide the first systematic framework for measuring and improving bystander privacy in audio LLMs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on6
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- SALMONN: Towards Generic Hearing Abilities for Large Language ModelsChangli Tang, Wenyi Yu, Guangzhi Sun, Xianzhao Chen et al.ICLR 2024 · 557 citations
- MMMU: A Massive Multi-Discipline Multimodal Understanding and Reasoning Benchmark for Expert AGIXiang Yue, Yuansheng Ni, Tianyu Zheng, Kai Zhang et al.CVPR 2024 · 213 citations
- SLMIA-SR: Speaker-Level Membership Inference Attacks against Speaker Recognition SystemsGuangke Chen, Yedi Zhang, Fu SongNDSS 2024
Related papers
- PrivLM-Bench: A Multi-level Privacy Evaluation Benchmark for Language ModelsHaoran Li, Dadi Guo, Donghao Li, Wei Fan et al.ACL 2024 · 9 citations
- VoxPrivacy: A Benchmark for Evaluating Interactional Privacy of Speech Language ModelsYuxiang Wang, HongYu Liu, Dekun Chen, Xueyao Zhang et al.ICLR 2026 · 5 citations
- PII-Bench: Evaluating Query-Aware Privacy Protection SystemsHao Shen, Zhouhong Gu, Haokai Hong, Weili Han et al.ACL 2026
- AIR-Bench: Benchmarking Large Audio-Language Models via Generative ComprehensionQian Yang, Jin Xu, Wenrui Liu, Yunfei Chu et al.ACL 2024
- Speech-Audio Compositional Attacks on Multimodal LLMs and Their Defense with SALMONN-GuardYudong Yang, Xuezhen Zhang, Zhifeng Han, Siyin Wang et al.ICML 2026 · 13 citations
