NeuroGenPoisoning: Neuron-Guided Attacks on Retrieval-Augmented Generation of LLM via Genetic Optimization of External Knowledge
Hanyu Zhu, Lance Fiondella, Jiawei Yuan, Kai Zeng, Long Jiao
摘要
Retrieval-Augmented Generation (RAG) empowers Large Language Models (LLMs) to dynamically integrate external knowledge during inference, improving their factual accuracy and adaptability. However, adversaries can inject poisoned external knowledge to override the model's internal memory. While existing attacks iteratively manipulate retrieval content or prompt structure of RAG, they largely ignore the model's internal representation dynamics and neuron-level sensitivities. The underlying mechanism of RAG poisoning has not been fully studied and the effect of knowledge conflict with strong parametric knowledge in RAG is not considered. In this work, we propose NeuroGenPoisoning, a novel attack framework that generates adversarial external knowledge in RAG guided by LLM internal neuron attribution and genetic optimization. Our method first identifies a set of Poison-Responsive Neurons whose activation strongly correlates with contextual poisoning knowledge. We then employ a genetic algorithm to evolve adversarial passages that maximally activate these neurons. Crucially, our framework enables massive-scale generation of effective poisoned RAG knowledge by identifying and reusing promising but initially unsuccessful external knowledge variants via observed attribution signals. At the same time, Poison-Responsive Neurons guided poisoning can effectively resolves knowledge conflict. Experimental results across models and datasets demonstrate consistently achieving high Population Overwrite Success Rate (POSR) of over 90% while preserving fluency. Empirical evidence shows that our method effectively resolves knowledge conflict. * Corresponding Author 39th Conference on Neural Information Processing Systems (NeurIPS 2025).
override the original knowledge of LLMs, leading to hallucinations, misinformation, or targeted disinformation [7,26,27,41,50].
Prior studies, such as PoisonedRAG [54], BadRAG [45], Pandora have demonstrated that LLMs can be manipulated by injecting carefully crafted contexts in RAG. However, these attacks typically rely on pre-defined misinformation templates or manually constructed adversarial passages [2,9,13,18,45,52,54], limiting the scalability and generality of their approach. Moreover, they do not explicitly model which internal components of the LLM are responsible for context reliance or knowledge conflicts. Recent research has identified a set of context-aware neurons [32], which are responsible for integrating external content into model predictions. If such neurons can be systematically activated by poisoned knowledge, then it is possible to override a model's parametric knowledge via a targeted manipulation of its internal decision pathway.
In this work, we propose NeuroGenPoisoning, a novel attack framework that leverages Poison-Responsive Neurons, neurons that are highly sensitive to external knowledge in RAG settings. Inspired by IRCAN [32], we identify Poison-Responsive Neurons via Integrated Gradients (IG) [33], and use their activation scores as optimization signals in genetic algorithms. We begin by prompting an LLM to generate misleading external knowledge passages containing a specified incorrect answer. These adversarial seeds mimic plausible sources while embedding targeted misinformation. From this initialization, we iteratively evolve the passages using a genetic algorithm guided by neuron attribution scores, which progressively amplify their influence on the model's output. By directly optimizing for internal attribution rather than surface-level cues, NeuroGenPoisoning crafts semantically coherent and stealthy poisoned knowledge capable of overriding the LLM's internal memory. Injected into the RAG pipeline, these optimized passages consistently induce hallucinations aligned with the adversary's target, even when the model has previously memorized the correct answer. Our experiments demonstrate that NeuroGenPoisoning can efficienlty launch massive RAG poisoning under large population and consistently achieves high Population Overwrite Success Rate (POSR) in multiple open-domain question answering (QA) datasets, including SQuAD 2.0 [30], TriviaQA [19], and WikiQA [46], and a variety of LLMs such as LLaMA-2-7b [36], Vicuna-7b/13b [8], and Gemma- 7b [35]. For example, on SQuAD 2.0, our method achieves a Population Overwrite Success Rate (POSR) of over 90% on LLaMA-2-7b-chat-hf, compared to an initial POSR of about 40% to 50%. Our method proves especially effective in knowledge conflict settings, in which the model has a strong internal memory of the correct answer. We observe that as genetic optimization progresses, the distribution of query-level POSR gradually shifts to higher, indicating that more queries achieve high POSR over time. Specially, the initialized external knowledge can lead to moderate POSR, with approximately 70% of queries exhibiting POSR between 40% and 50%. However, the success remains inconsistent, with a minority of queries succeed completely (
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Visual Inception: Compromising Long-term Planning in Agentic Recommenders via Multimodal Memory PoisoningJiachen QianACL 2026 · 被引用 3 次
- InceptionRAG: Stealthy Poisoning Attack Against Retrieval-Augmented GenerationJiachang Zhang, Min Chen, Xiao Ren, Zhenyong Zhang 等CCS 2026
它引用的顶会 Paper24
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-ReflectionAkari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil 等ICLR 2024 · 被引用 1,798 次
- Improving Language Models by Retrieving from Trillions of TokensSebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai 等ICML 2022 · 被引用 1,629 次
- Tree of Attacks: Jailbreaking Black-Box LLMs AutomaticallyAnay Mehrotra, Manolis Zampetakis, Paul Kassianik, Blaine Nelson 等NeurIPS 2024 · 被引用 835 次
- AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language ModelsXiaogeng Liu, Nan Xu, Muhao Chen, Chaowei XiaoICLR 2024 · 被引用 722 次
相关 Paper
- MM-PoisonRAG: Disrupting Multimodal RAG with Local and Global Knowledge Poisoning AttacksHyeonjeong Ha, Qiusi Zhan, Jeonghwan Kim, Dimitrios Bralios 等ACL 2026 · 被引用 1 次
- KEPo: Knowledge Evolution Poison on Graph-based Retrieval-Augmented GenerationQizhi Chen, Chao Qi, Yihong Huang, Muquan Li 等WWW 2026
- Joint-GCG: Unified Gradient-Based Poisoning Attacks on Retrieval-Augmented Generation SystemsHaowei Wang, Rupeng Zhang, Junjie Wang, Mingyang Li 等AAAI 2026 · 被引用 3 次
- On the Vulnerability of Applying Retrieval-Augmented Generation within Knowledge-Intensive Application DomainsXun Xian, Ganghua Wang, Xuan Bi, Rui Zhang 等ICML 2025
- PR-Attack: Coordinated Prompt-RAG Attacks on Retrieval-Augmented Generation in Large Language Models via Bilevel OptimizationYang Jiao, Xiaodong Wang, Kai YangSIGIR 2025 · 被引用 6 次
