ALGEN: Few-shot Inversion Attacks on Textual Embeddings via Cross-Model Alignment and Generation
Yiyi Chen, Qiongkai Xu, Johannes Bjerva
Abstract
With the growing popularity of Large Language Models (LLMs) and vector databases, private textual data is increasingly processed and stored as numerical embeddings. However, recent studies have proven that such embeddings are vulnerable to inversion attacks, where original text is reconstructed to reveal sensitive information. Previous research has largely assumed access to millions of sentences to train attack models, e.g., through data leakage or nearly unrestricted API access. With our method, a single data point is sufficient for a partially successful inversion attack. With as little as 1k data samples, performance reaches an optimum across a range of black-box encoders, without training on leaked data. We present a Few-shot Textual Embedding Inversion Attack using Cross-Model ALignment and GENeration (ALGEN), by aligning victim embeddings to the attack space and using a generative model to reconstruct text. We find that ALGEN attacks can be effectively transferred across domains and languages, revealing key information. We further examine a variety of defense mechanisms against ALGEN, and find that none are effective, highlighting the vulnerabilities posed by inversion attacks. By significantly lowering the cost of inversion and proving that embedding spaces can be aligned through one-step optimization, we establish a new textual embedding inversion paradigm with broader applications for embedding alignment in NLP. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f07b42d7-aedd-42ab-b49b-208d95c5c216Cited by top-tier papers2
- Towards Whole-corpus Reconstruction of Heterogeneous RAG Knowledge BasesPeiru Yang, Yi Luo, Zhenfeng Gao, Tong Ju et al.ICML 2026
- FedAugment: Table Augmentation Search over Decentralized Data RepositoriesLennart Behme, Emil Badura, Leonard Geißler, Matthias Böhm et al.VLDB 2026
Builds on7
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Multilingual Alignment of Contextual Word RepresentationsSteven Cao, Nikita Kitaev, Dan KleinICLR 2020 · 211 citations
- Information Leakage in Embedding ModelsCongzheng Song, Ananth RaghunathanCCS 2020 · 200 citations
- Text Embeddings Reveal (Almost) As Much As TextJohn X. Morris, Volodymyr Kuleshov, Vitaly Shmatikov, Alexander M. RushEMNLP 2023 · 60 citations
- Normalization of Language Embeddings for Cross-Lingual AlignmentPrince Osei Aboagye, Yan Zheng, Chin-Chia Michael Yeh, Junpeng Wang et al.ICLR 2022 · 12 citations
Related papers
- Text Embedding Inversion Security for Multilingual Language ModelsYiyi Chen, Heather C. Lent, Johannes BjervaACL 2024 · 10 citations
- Transferable Embedding Inversion Attack: Uncovering Privacy Risks in Text Embeddings without Model QueriesYu-Hsiang Huang, Yu-Che Tsai, Hsiang Hsiao, Hong-Yi Lin et al.ACL 2024 · 5 citations
- Black-Box Embedding Inversion Attack on Vector DatabasesLichao Sun, Yuncheng Wu, Haichao Sha, Xinjian Luo et al.KDD 2026
- Overcoming the Retrieval Barrier: Indirect Prompt Injection in the Wild for LLM SystemsHongyan Chang, Ergute Bao, Xinjian Luo, Ting YuUSENIX Security 2026 · 24 citations
- Depth Gives a False Sense of Privacy: LLM Internal States InversionTian Dong, Yan Meng, Shaofeng Li, Guoxing Chen et al.USENIX Security 2025
