Humpty Dumpty: Controlling Word Meanings via Corpus Poisoning
Roei Schuster, Tal Schuster, Yoav Meri, Vitaly Shmatikov
Abstract
Word embeddings, i.e., low-dimensional vector representations such as GloVe and SGNS, encode word "meaning" in the sense that distances between words’ vectors correspond to their semantic proximity. This enables transfer learning of semantics for a variety of natural language processing tasks.Word embeddings are typically trained on large public corpora such as Wikipedia or Twitter. We demonstrate that an attacker who can modify the corpus on which the embedding is trained can control the "meaning" of new and existing words by changing their locations in the embedding space. We develop an explicit expression over corpus features that serves as a proxy for distance between words and establish a causative relationship between its values and embedding distances. We then show how to use this relationship for two adversarial objectives: (1) make a word a top-ranked neighbor of another word, and (2) move a word from one semantic cluster to another.An attack on the embedding can affect diverse downstream tasks, demonstrating for the first time the power of data poisoning in transfer learning scenarios. We use this attack to manipulate query expansion in information retrieval systems such as resume search, make certain names more or less visible to named entity recognition models, and cause new words to be translated to a particular target word regardless of the language. Finally, we show how the attacker can generate linguistically likely corpus modifications, thus fooling defenses that attempt to filter implausible sentences from the corpus using a language model.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers11
- Weight Poisoning Attacks on Pretrained ModelsKeita Kurita, Paul Michel, Graham NeubigACL 2020 · 312 citations
- You Autocomplete Me: Poisoning Vulnerabilities in Neural Code CompletionRoei Schuster, Congzheng Song, Eran Tromer, Vitaly ShmatikovUSENIX Security 2021 · 199 citations
- Universal Jailbreak Backdoors from Poisoned Human FeedbackJavier Rando, Florian TramèrICLR 2024 · 124 citations
- Spinning Language Models: Risks of Propaganda-As-A-Service and CountermeasuresEugene Bagdasaryan, Vitaly ShmatikovS&P 2022 · 94 citations
- Reverse Attack: Black-box Attacks on Collaborative RecommendationYihe Zhang, Xu Yuan, Jin Li, Jiadong Lou et al.CCS 2021 · 24 citations
Builds on1
Related papers
- Unsupervised Corpus Poisoning Attacks in Continuous Space for Dense RetrievalYongkang Li, Panagiotis Eustratiadis, Simon Lupart, Evangelos KanoulasSIGIR 2025 · 3 citations
- Information Leakage in Embedding ModelsCongzheng Song, Ananth RaghunathanCCS 2020 · 200 citations
- Transferable Embedding Inversion Attack: Uncovering Privacy Risks in Text Embeddings without Model QueriesYu-Hsiang Huang, Yu-Che Tsai, Hsiang Hsiao, Hong-Yi Lin et al.ACL 2024 · 5 citations
- Thieves on Sesame Street! Model Extraction of BERT-based APIsKalpesh Krishna, Gaurav Singh Tomar, Ankur P. Parikh, Nicolas Papernot et al.ICLR 2020 · 244 citations
- ``Someone Hid It!'': Query-Agnostic Black-Box Attacks on LLM-Based RetrievalJiate Li, Defu Cao, Li Li, Wei Yang et al.ICML 2026 · 4 citations
