Black-Box Embedding Inversion Attack on Vector Databases
Lichao Sun, Yuncheng Wu, Haichao Sha, Xinjian Luo, Mingyang Yi, Meihui Zhang, Cuiping Li, Hong Chen
摘要
Vector databases that index and serve dense embeddings have become central to modern data science applications. Embeddings are often regarded as privacy-preserving surrogates for raw data, motivating practitioners to outsource vector databases to third-party services for scalability. However, recent studies show that even text embeddings alone can leak sensitive information, raising serious privacy concerns. Existing attacks on image embeddings, meanwhile, typically assume access to model architecture or parameters, which does not hold in outsourced settings. In this paper, we propose a novel black-box image embedding inversion attack that reconstructs high-fidelity images using only query access to the embedding model or API. Our approach leverages an in-distribution auxiliary dataset to train a conditional diffusion model, capturing domain-aligned knowledge of the data owner's private images. We introduce an embedding-guided cross-attention mechanism, where image embeddings serve as conditional signals to steer the generation process. To improve efficiency, we perform diffusion in the latent space of a pretrained VQGAN with deterministic decoding, which reduces the computational cost while preserving both structural and perceptual fidelity in reconstructed images. Extensive experiments on three real-world datasets and three popular embedding models demonstrate that our approach significantly outperforms four state-of-the-art baselines. These findings reveal that image embeddings can expose sensitive visual information, highlighting the need for stronger privacy protections in outsourced vector databases.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Black-box Membership Inference Attacks against Fine-tuned Diffusion ModelsYan Pang, Tianhao WangNDSS 2025
- Enhanced Privacy Leakage from Noise-Perturbed Gradients via Gradient-Guided Conditional Diffusion ModelsJiayang Meng, Tao Huang, Hong Chen, Chen Hou 等AAAI 2026 · 被引用 1 次
- Reinforcement Learning-Based Black-Box Model Inversion AttacksGyojin Han, Jaehyun Choi, Haeil Lee, Junmo KimCVPR 2023
- ALGEN: Few-shot Inversion Attacks on Textual Embeddings via Cross-Model Alignment and GenerationYiyi Chen, Qiongkai Xu, Johannes BjervaACL 2025
- Transferable Embedding Inversion Attack: Uncovering Privacy Risks in Text Embeddings without Model QueriesYu-Hsiang Huang, Yu-Che Tsai, Hsiang Hsiao, Hong-Yi Lin 等ACL 2024 · 被引用 5 次
