Lune

S&P2026顶会

The Person Behind the Sound: Demystifying Audio Private Attribute Profiling Via Multimodal Large Language Models

Lixu Wang, Kaixiang Yao, Xinfeng Li, Dong Yang, Haoyao Li, Xiaofeng Wang, Wei Dong

2026年份

摘要

Our research uncovers a severe privacy risk associated with multimodal large language models (MLLMs): the ability to infer sensitive personal attributes from audio data, which we call audio private attribute profiling. This capability poses a significant threat, as audio can be covertly captured using simple tools. Moreover, compared to images and texts, audio carries unique characteristics, such as tone and pitch, which can be exploited for more detailed attribute profiling. The first major barrier to understanding this threat is the lack of benchmark datasets with profile-level sensitive attribute annotations. Collecting audio data with attribute labels from real-world volunteers is impractical due to legal, ethical, and compliance concerns. To address this challenge, we introduce AP2\text{AP}^{2}, a well-crafted audio benchmark dataset constructed using public sources and recent TV dramas. On AP2A P^{2}, we examine two baseline avenues of profiling sensitive attributes: (1) converting audio to text and applying LLMs, and (2) directly using audiolanguage models (ALMs). We found that the former suffers from information loss during transcription, while the latter lacks sufficient reasoning capability. To overcome these limitations, we propose Gifts, a hybrid framework in which an LLM guides, forensically reviews, and consolidates inferences made by an ALM. Gifts mitigates information loss by letting the ALM lead the inference, while the LLM enhances inference accuracy and validity through three phases: guidance, review, and consolidation. Extensive experiments and human evaluations of 50\mathbf{5 0} participants (18-30 years) show that Gifts outperforms the MLLM-based baselines, real humans, and traditional inference methods in profiling sensitive attributes, while also being robust under various types of noise. We further study defense strategies at both the model and data levels. Our work demonstrates the feasibility of audio privacy leakage caused by MLLMs, highlights the urgent need for effective defenses, and provides resources to support future research.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖