CapRecover: A Cross-Modality Feature Inversion Attack Framework on Vision Language Models
Kedong Xiu, Sai Qian Zhang
摘要
As Vision-Language Models (VLMs) become increasingly integrated into user-facing applications, they are often deployed in split DNN configurations, where the visual encoder (e.g., ResNet or ViT) runs on user-side devices and only intermediate features are transmitted to the cloud for downstream processing. While this setup reduces communication overhead, the intermediate data features containing sensitive information can also expose users to privacy risks. Prior work has attempted to reconstruct images from these features to infer semantics, but such approaches often produce blurry images that obscure semantic details. In contrast, the potential to directly recover high-level semantic content -such as image labels or captions -via a cross-modality inversion attack remains largely unexplored. To address this gap, we propose CapRecover, a general cross-modality feature inversion framework that directly decodes semantic information from intermediate features without requiring image reconstruction. Additionally, CapRecover can be used to reverse engineer traditional neural networks for computer vision tasks, such as ViT, ResNet, and others.
We evaluate CapRecover across multiple widely used datasets and victim models. Our results demonstrate that CapRecover can accurately recover both image labels and captions without reconstructing a single pixel. Specifically, it achieves up to 92.71% Top-1 accuracy on the CIFAR-10 dataset for label recovery, and generates fluent and relevant captions from ResNet50's intermediate features on COCO2017 dataset, with ROUGE-L scores up to 0.52. Furthermore, an in-depth analysis of ResNet-based models reveals that deeper convolutional layers encode significantly more semantic information, whereas shallow layers contribute minimally to semantic leakage. Furthermore, we propose a straightforward and effective protection approach that adds random noise to the intermediate image features at each middle layer and subsequently removes the noise in the following layer. Our experiments indicate that this approach effectively prevents information leakage without additional training costs. Our code is available here.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le 等ICCV 2019 · 被引用 9,163 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- EfficientFormer: Vision Transformers at MobileNet SpeedYanyu Li, Geng Yuan, Yang Wen, Ju Hu 等NeurIPS 2022 · 被引用 742 次
相关 Paper
- What Your Features Reveal: Data-Efficient Black-Box Feature Inversion Attack for Split DNNsZhihan Ren, Lijun He, Jiaxi Liang, Xinzhu Fu 等CVPR 2026 · 被引用 2 次
- Do Vision-Language Models Leak What They Learn? Adaptive Token-Weighted Model Inversion AttacksNgoc-Bao Nguyen, Sy-Tuyen Ho, Koh Jun Hao, Ngai-Man CheungCVPR 2026 · 被引用 2 次
- Gradient Inversion of Multimodal ModelsOmri Ben Hemo, Alon Zolfi, Oryan Yehezkel, Omer Hofman 等ICML 2025
- Reversible Privacy Preserving on Vision-Language Models via Adversarial Multimodal KeyPeng Ying, Zhongnian Li, Meng Wei, Xinzheng XuACM MM 2025
- Prompt Inference Attack on Distributed Large Language Model Inference FrameworksXinjian Luo, Ting Yu, Xiaokui XiaoCCS 2025
