Steering Representations, Safeguarding Privacy: A Cross-Modal Privacy Protection Method for Generative AI
Jie Zhang, Chenxu Niu, Zhefeng Nan, Yangyan Xu, Jinta Weng
摘要
Privacy concerns have long been a critical issue in AI models. With the rapid advancement of generative AI, the privacy awareness of models has drawn attention, raising new challenges for privacy protection that is independent of data and tasks. This paper introduces a novel framework for enhancing privacy protection through directional steering in representation space, which seamlessly integrates with both language and vision-language models. Specifically, we first construct a comprehensive privacy-related dataset based on the Solove taxonomy of privacy. Then, we leverage this dataset to enhance model privacy awareness in the representation space, steering the model to protect privacy during inference. Experiments on 12 models validate the effectiveness and generalization of our method. Moreover, we demonstrate the transferability of privacy-enhanced representations between same-source large language models (LLMs) and vision-language models (VLMs), offering a scalable solution for privacy protection in frontier AI models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- Inference-Time Intervention: Eliciting Truthful Answers from a Language ModelKenneth Li, Oam Patel, Fernanda B. Viégas, Hanspeter Pfister 等NeurIPS 2023 · 被引用 1,549 次
- Self-Instruct: Aligning Language Models with Self-Generated InstructionsYizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu 等ACL 2023 · 被引用 540 次
- The Linear Representation Hypothesis and the Geometry of Large Language ModelsKiho Park, Yo Joong Choe, Victor VeitchICML 2024 · 被引用 461 次
- Privacy Risks of General-Purpose Language ModelsXudong Pan, Mi Zhang, Shouling Ji, Min YangS&P 2020 · 被引用 291 次
- ProPILE: Probing Privacy Leakage in Large Language ModelsSiwon Kim, Sangdoo Yun, Hwaran Lee, Martin Gubri 等NeurIPS 2023 · 被引用 229 次
相关 Paper
- The Eye of Sherlock Holmes: Uncovering User Private Attribute Profiling via Vision-Language Model Agentic FrameworkFeiran Liu, Yuzhe Zhang, Xinyi Huang, Yinan Peng 等ACM MM 2025 · 被引用 4 次
- PrivSV: Differentially Private Steering Vector for Large Language ModelsHaocheng Yang, Xiang Cheng, Chenhao Sun, Pengfei Zhang 等AAAI 2026
- When Privacy Meets Recovery: The Overlooked Half of Surrogate-Driven Privacy Preservation for MLLM EditingSiyuan Xu, Yibing Liu, Peilin Chen, Yung-Hui Li 等AAAI 2026
- Boundary Probing for Input Privacy Protection when Using LMM ServicesXiaofei Hui, Haoxuan Qu, Ping Hu, Hossein Rahmani 等ICCV 2025
- Private Attribute Inference from Images with Vision-Language ModelsBatuhan Tömekçe, Mark Vero, Robin Staab, Martin T. VechevNeurIPS 2024 · 被引用 54 次
