The Person Behind the Sound: Demystifying Audio Private Attribute Profiling Via Multimodal Large Language Models
Lixu Wang, Kaixiang Yao, Xinfeng Li, Dong Yang, Haoyao Li, Xiaofeng Wang, Wei Dong
Abstract
Our research uncovers a severe privacy risk associated with multimodal large language models (MLLMs): the ability to infer sensitive personal attributes from audio data, which we call audio private attribute profiling. This capability poses a significant threat, as audio can be covertly captured using simple tools. Moreover, compared to images and texts, audio carries unique characteristics, such as tone and pitch, which can be exploited for more detailed attribute profiling. The first major barrier to understanding this threat is the lack of benchmark datasets with profile-level sensitive attribute annotations. Collecting audio data with attribute labels from real-world volunteers is impractical due to legal, ethical, and compliance concerns. To address this challenge, we introduce , a well-crafted audio benchmark dataset constructed using public sources and recent TV dramas. On , we examine two baseline avenues of profiling sensitive attributes: (1) converting audio to text and applying LLMs, and (2) directly using audiolanguage models (ALMs). We found that the former suffers from information loss during transcription, while the latter lacks sufficient reasoning capability. To overcome these limitations, we propose Gifts, a hybrid framework in which an LLM guides, forensically reviews, and consolidates inferences made by an ALM. Gifts mitigates information loss by letting the ALM lead the inference, while the LLM enhances inference accuracy and validity through three phases: guidance, review, and consolidation. Extensive experiments and human evaluations of participants (18-30 years) show that Gifts outperforms the MLLM-based baselines, real humans, and traditional inference methods in profiling sensitive attributes, while also being robust under various types of noise. We further study defense strategies at both the model and data levels. Our work demonstrates the feasibility of audio privacy leakage caused by MLLMs, highlights the urgent need for effective defenses, and provides resources to support future research.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 65c9ad67-1c24-4e29-95e5-31dfe3f77d9bRelated papers
- The Eye of Sherlock Holmes: Uncovering User Private Attribute Profiling via Vision-Language Model Agentic FrameworkFeiran Liu, Yuzhe Zhang, Xinyi Huang, Yinan Peng et al.ACM MM 2025 · 4 citations
- AudioTrust: Benchmarking The Multifaceted Trustworthiness of Audio Large Language ModelsKai Li, Can Shen, Yile Liu, Jirui Han et al.ICLR 2026 · 17 citations
- MultiPriv: Benchmarking Individual-Level Privacy Reasoning in Vision-Language ModelsXiongtao Sun, HUI LI, Jiaming Zhang, Yujie Yang et al.ICML 2026 · 3 citations
- Private Attribute Inference from Images with Vision-Language ModelsBatuhan Tömekçe, Mark Vero, Robin Staab, Martin T. VechevNeurIPS 2024 · 54 citations
- PIG: Privacy Jailbreak Attack on LLMs via Gradient-based Iterative In-Context OptimizationYidan Wang, Yanan Cao, Yubing Ren, Fang Fang et al.ACL 2025 · 12 citations
