Black-Box Embedding Inversion Attack on Vector Databases
Lichao Sun, Yuncheng Wu, Haichao Sha, Xinjian Luo, Mingyang Yi, Meihui Zhang, Cuiping Li, Hong Chen
Abstract
Vector databases that index and serve dense embeddings have become central to modern data science applications. Embeddings are often regarded as privacy-preserving surrogates for raw data, motivating practitioners to outsource vector databases to third-party services for scalability. However, recent studies show that even text embeddings alone can leak sensitive information, raising serious privacy concerns. Existing attacks on image embeddings, meanwhile, typically assume access to model architecture or parameters, which does not hold in outsourced settings. In this paper, we propose a novel black-box image embedding inversion attack that reconstructs high-fidelity images using only query access to the embedding model or API. Our approach leverages an in-distribution auxiliary dataset to train a conditional diffusion model, capturing domain-aligned knowledge of the data owner's private images. We introduce an embedding-guided cross-attention mechanism, where image embeddings serve as conditional signals to steer the generation process. To improve efficiency, we perform diffusion in the latent space of a pretrained VQGAN with deterministic decoding, which reduces the computational cost while preserving both structural and perceptual fidelity in reconstructed images. Extensive experiments on three real-world datasets and three popular embedding models demonstrate that our approach significantly outperforms four state-of-the-art baselines. These findings reveal that image embeddings can expose sensitive visual information, highlighting the need for stronger privacy protections in outsourced vector databases.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 47a269e0-1609-4e2c-8e57-9e03ef201270Related papers
- Black-box Membership Inference Attacks against Fine-tuned Diffusion ModelsYan Pang, Tianhao WangNDSS 2025
- Enhanced Privacy Leakage from Noise-Perturbed Gradients via Gradient-Guided Conditional Diffusion ModelsJiayang Meng, Tao Huang, Hong Chen, Chen Hou et al.AAAI 2026 · 1 citation
- Reinforcement Learning-Based Black-Box Model Inversion AttacksGyojin Han, Jaehyun Choi, Haeil Lee, Junmo KimCVPR 2023
- ALGEN: Few-shot Inversion Attacks on Textual Embeddings via Cross-Model Alignment and GenerationYiyi Chen, Qiongkai Xu, Johannes BjervaACL 2025
- Transferable Embedding Inversion Attack: Uncovering Privacy Risks in Text Embeddings without Model QueriesYu-Hsiang Huang, Yu-Che Tsai, Hsiang Hsiao, Hong-Yi Lin et al.ACL 2024 · 5 citations
