The Person Behind the Sound: Demystifying Audio Private Attribute Profiling Via Multimodal Large Language Models
Lixu Wang, Kaixiang Yao, Xinfeng Li, Dong Yang, Haoyao Li, Xiaofeng Wang, Wei Dong
摘要
Our research uncovers a severe privacy risk associated with multimodal large language models (MLLMs): the ability to infer sensitive personal attributes from audio data, which we call audio private attribute profiling. This capability poses a significant threat, as audio can be covertly captured using simple tools. Moreover, compared to images and texts, audio carries unique characteristics, such as tone and pitch, which can be exploited for more detailed attribute profiling. The first major barrier to understanding this threat is the lack of benchmark datasets with profile-level sensitive attribute annotations. Collecting audio data with attribute labels from real-world volunteers is impractical due to legal, ethical, and compliance concerns. To address this challenge, we introduce , a well-crafted audio benchmark dataset constructed using public sources and recent TV dramas. On , we examine two baseline avenues of profiling sensitive attributes: (1) converting audio to text and applying LLMs, and (2) directly using audiolanguage models (ALMs). We found that the former suffers from information loss during transcription, while the latter lacks sufficient reasoning capability. To overcome these limitations, we propose Gifts, a hybrid framework in which an LLM guides, forensically reviews, and consolidates inferences made by an ALM. Gifts mitigates information loss by letting the ALM lead the inference, while the LLM enhances inference accuracy and validity through three phases: guidance, review, and consolidation. Extensive experiments and human evaluations of participants (18-30 years) show that Gifts outperforms the MLLM-based baselines, real humans, and traditional inference methods in profiling sensitive attributes, while also being robust under various types of noise. We further study defense strategies at both the model and data levels. Our work demonstrates the feasibility of audio privacy leakage caused by MLLMs, highlights the urgent need for effective defenses, and provides resources to support future research.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- The Eye of Sherlock Holmes: Uncovering User Private Attribute Profiling via Vision-Language Model Agentic FrameworkFeiran Liu, Yuzhe Zhang, Xinyi Huang, Yinan Peng 等ACM MM 2025 · 被引用 4 次
- AudioTrust: Benchmarking The Multifaceted Trustworthiness of Audio Large Language ModelsKai Li, Can Shen, Yile Liu, Jirui Han 等ICLR 2026 · 被引用 17 次
- MultiPriv: Benchmarking Individual-Level Privacy Reasoning in Vision-Language ModelsXiongtao Sun, HUI LI, Jiaming Zhang, Yujie Yang 等ICML 2026 · 被引用 3 次
- Private Attribute Inference from Images with Vision-Language ModelsBatuhan Tömekçe, Mark Vero, Robin Staab, Martin T. VechevNeurIPS 2024 · 被引用 54 次
- PIG: Privacy Jailbreak Attack on LLMs via Gradient-based Iterative In-Context OptimizationYidan Wang, Yanan Cao, Yubing Ren, Fang Fang 等ACL 2025 · 被引用 12 次
