Reversible Privacy Preserving on Vision-Language Models via Adversarial Multimodal Key
Peng Ying, Zhongnian Li, Meng Wei, Xinzheng Xu
Abstract
Vision-Language Models (VLMs) such as GPT-4V and LLaVA have demonstrated impressive capabilities in multimodal understanding and generation. Unfortunately, their ability to infer sensitive information from visual content raises serious privacy concerns, especially when the images containing personal information. Existing solutions either rely on static alignment mechanisms, such as task-specific prompt turning, which are vulnerable to adversarial prompts, or irreversible redaction methods that permanently destroy content utility for legitimate users. To address these limitations, we propose a reversible privacy-preserving framework on VLMs via Adversarial Multimodal Key (AMK). Specifically, AMK embeds a learnable adversarial image key into mosaic-obscured images and generates a corresponding text key through multimodal contrastive learning model. The image key is optimized by gradient-based supervision from white-box VLMs, and the text key is implicitly derived from the image content to avoid exposure during transmission. These keys enable authorized users to restore sensitive information through VLMs, while unauthorized queries are explicitly rejected through a refusal response mechanism. Our experiments across various privacy scenarios show that the proposed method effectively restores redacted content with correct keys and prevents unauthorized disclosure, offering a practical solution for privacy protection in multimodal systems.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- When Privacy Meets Recovery: The Overlooked Half of Surrogate-Driven Privacy Preservation for MLLM EditingSiyuan Xu, Yibing Liu, Peilin Chen, Yung-Hui Li et al.AAAI 2026
- Boundary Probing for Input Privacy Protection when Using LMM ServicesXiaofei Hui, Haoxuan Qu, Ping Hu, Hossein Rahmani et al.ICCV 2025
- GeoShield: Safeguarding Geolocation Privacy from Vision-Language Models via Adversarial PerturbationsXinwei Liu, Xiaojun Jia, Yuan Xun, Simeng Qin et al.AAAI 2026 · 2 citations
- Leave My Images Alone: Preventing Multi-Modal Large Language Models from Analyzing Images via Visual Prompt InjectionZedian Shao, Hongbin Liu, Yuepeng Hu, Neil Zhenqiang GongACL 2026
- VLSBench: Unveiling Visual Leakage in Multimodal SafetyXuhao Hu, Dongrui Liu, Hao Li, Xuanjing Huang et al.ACL 2025
