ALGEN: Few-shot Inversion Attacks on Textual Embeddings via Cross-Model Alignment and Generation
Yiyi Chen, Qiongkai Xu, Johannes Bjerva
摘要
With the growing popularity of Large Language Models (LLMs) and vector databases, private textual data is increasingly processed and stored as numerical embeddings. However, recent studies have proven that such embeddings are vulnerable to inversion attacks, where original text is reconstructed to reveal sensitive information. Previous research has largely assumed access to millions of sentences to train attack models, e.g., through data leakage or nearly unrestricted API access. With our method, a single data point is sufficient for a partially successful inversion attack. With as little as 1k data samples, performance reaches an optimum across a range of black-box encoders, without training on leaked data. We present a Few-shot Textual Embedding Inversion Attack using Cross-Model ALignment and GENeration (ALGEN), by aligning victim embeddings to the attack space and using a generative model to reconstruct text. We find that ALGEN attacks can be effectively transferred across domains and languages, revealing key information. We further examine a variety of defense mechanisms against ALGEN, and find that none are effective, highlighting the vulnerabilities posed by inversion attacks. By significantly lowering the cost of inversion and proving that embedding spaces can be aligned through one-step optimization, we establish a new textual embedding inversion paradigm with broader applications for embedding alignment in NLP. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Towards Whole-corpus Reconstruction of Heterogeneous RAG Knowledge BasesPeiru Yang, Yi Luo, Zhenfeng Gao, Tong Ju 等ICML 2026
- FedAugment: Table Augmentation Search over Decentralized Data RepositoriesLennart Behme, Emil Badura, Leonard Geißler, Matthias Böhm 等VLDB 2026
它引用的顶会 Paper7
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Multilingual Alignment of Contextual Word RepresentationsSteven Cao, Nikita Kitaev, Dan KleinICLR 2020 · 被引用 211 次
- Information Leakage in Embedding ModelsCongzheng Song, Ananth RaghunathanCCS 2020 · 被引用 200 次
- Text Embeddings Reveal (Almost) As Much As TextJohn X. Morris, Volodymyr Kuleshov, Vitaly Shmatikov, Alexander M. RushEMNLP 2023 · 被引用 60 次
- Normalization of Language Embeddings for Cross-Lingual AlignmentPrince Osei Aboagye, Yan Zheng, Chin-Chia Michael Yeh, Junpeng Wang 等ICLR 2022 · 被引用 12 次
相关 Paper
- Text Embedding Inversion Security for Multilingual Language ModelsYiyi Chen, Heather C. Lent, Johannes BjervaACL 2024 · 被引用 10 次
- Transferable Embedding Inversion Attack: Uncovering Privacy Risks in Text Embeddings without Model QueriesYu-Hsiang Huang, Yu-Che Tsai, Hsiang Hsiao, Hong-Yi Lin 等ACL 2024 · 被引用 5 次
- Black-Box Embedding Inversion Attack on Vector DatabasesLichao Sun, Yuncheng Wu, Haichao Sha, Xinjian Luo 等KDD 2026
- Overcoming the Retrieval Barrier: Indirect Prompt Injection in the Wild for LLM SystemsHongyan Chang, Ergute Bao, Xinjian Luo, Ting YuUSENIX Security 2026 · 被引用 24 次
- Depth Gives a False Sense of Privacy: LLM Internal States InversionTian Dong, Yan Meng, Shaofeng Li, Guoxing Chen 等USENIX Security 2025
