NeuroGenPoisoning: Neuron-Guided Attacks on Retrieval-Augmented Generation of LLM via Genetic Optimization of External Knowledge
Hanyu Zhu, Lance Fiondella, Jiawei Yuan, Kai Zeng, Long Jiao
Abstract
Retrieval-Augmented Generation (RAG) empowers Large Language Models (LLMs) to dynamically integrate external knowledge during inference, improving their factual accuracy and adaptability. However, adversaries can inject poisoned external knowledge to override the model's internal memory. While existing attacks iteratively manipulate retrieval content or prompt structure of RAG, they largely ignore the model's internal representation dynamics and neuron-level sensitivities. The underlying mechanism of RAG poisoning has not been fully studied and the effect of knowledge conflict with strong parametric knowledge in RAG is not considered. In this work, we propose NeuroGenPoisoning, a novel attack framework that generates adversarial external knowledge in RAG guided by LLM internal neuron attribution and genetic optimization. Our method first identifies a set of Poison-Responsive Neurons whose activation strongly correlates with contextual poisoning knowledge. We then employ a genetic algorithm to evolve adversarial passages that maximally activate these neurons. Crucially, our framework enables massive-scale generation of effective poisoned RAG knowledge by identifying and reusing promising but initially unsuccessful external knowledge variants via observed attribution signals. At the same time, Poison-Responsive Neurons guided poisoning can effectively resolves knowledge conflict. Experimental results across models and datasets demonstrate consistently achieving high Population Overwrite Success Rate (POSR) of over 90% while preserving fluency. Empirical evidence shows that our method effectively resolves knowledge conflict. * Corresponding Author 39th Conference on Neural Information Processing Systems (NeurIPS 2025).
override the original knowledge of LLMs, leading to hallucinations, misinformation, or targeted disinformation [7,26,27,41,50].
Prior studies, such as PoisonedRAG [54], BadRAG [45], Pandora have demonstrated that LLMs can be manipulated by injecting carefully crafted contexts in RAG. However, these attacks typically rely on pre-defined misinformation templates or manually constructed adversarial passages [2,9,13,18,45,52,54], limiting the scalability and generality of their approach. Moreover, they do not explicitly model which internal components of the LLM are responsible for context reliance or knowledge conflicts. Recent research has identified a set of context-aware neurons [32], which are responsible for integrating external content into model predictions. If such neurons can be systematically activated by poisoned knowledge, then it is possible to override a model's parametric knowledge via a targeted manipulation of its internal decision pathway.
In this work, we propose NeuroGenPoisoning, a novel attack framework that leverages Poison-Responsive Neurons, neurons that are highly sensitive to external knowledge in RAG settings. Inspired by IRCAN [32], we identify Poison-Responsive Neurons via Integrated Gradients (IG) [33], and use their activation scores as optimization signals in genetic algorithms. We begin by prompting an LLM to generate misleading external knowledge passages containing a specified incorrect answer. These adversarial seeds mimic plausible sources while embedding targeted misinformation. From this initialization, we iteratively evolve the passages using a genetic algorithm guided by neuron attribution scores, which progressively amplify their influence on the model's output. By directly optimizing for internal attribution rather than surface-level cues, NeuroGenPoisoning crafts semantically coherent and stealthy poisoned knowledge capable of overriding the LLM's internal memory. Injected into the RAG pipeline, these optimized passages consistently induce hallucinations aligned with the adversary's target, even when the model has previously memorized the correct answer. Our experiments demonstrate that NeuroGenPoisoning can efficienlty launch massive RAG poisoning under large population and consistently achieves high Population Overwrite Success Rate (POSR) in multiple open-domain question answering (QA) datasets, including SQuAD 2.0 [30], TriviaQA [19], and WikiQA [46], and a variety of LLMs such as LLaMA-2-7b [36], Vicuna-7b/13b [8], and Gemma- 7b [35]. For example, on SQuAD 2.0, our method achieves a Population Overwrite Success Rate (POSR) of over 90% on LLaMA-2-7b-chat-hf, compared to an initial POSR of about 40% to 50%. Our method proves especially effective in knowledge conflict settings, in which the model has a strong internal memory of the correct answer. We observe that as genetic optimization progresses, the distribution of query-level POSR gradually shifts to higher, indicating that more queries achieve high POSR over time. Specially, the initialized external knowledge can lead to moderate POSR, with approximately 70% of queries exhibiting POSR between 40% and 50%. However, the success remains inconsistent, with a minority of queries succeed completely (
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3870c312-db71-46b8-9e9b-b7cf6506f501Cited by top-tier papers2
- Visual Inception: Compromising Long-term Planning in Agentic Recommenders via Multimodal Memory PoisoningJiachen QianACL 2026 · 3 citations
- InceptionRAG: Stealthy Poisoning Attack Against Retrieval-Augmented GenerationJiachang Zhang, Min Chen, Xiao Ren, Zhenyong Zhang et al.CCS 2026
Builds on24
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-ReflectionAkari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil et al.ICLR 2024 · 1,798 citations
- Improving Language Models by Retrieving from Trillions of TokensSebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai et al.ICML 2022 · 1,629 citations
- Tree of Attacks: Jailbreaking Black-Box LLMs AutomaticallyAnay Mehrotra, Manolis Zampetakis, Paul Kassianik, Blaine Nelson et al.NeurIPS 2024 · 835 citations
- AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language ModelsXiaogeng Liu, Nan Xu, Muhao Chen, Chaowei XiaoICLR 2024 · 722 citations
Related papers
- MM-PoisonRAG: Disrupting Multimodal RAG with Local and Global Knowledge Poisoning AttacksHyeonjeong Ha, Qiusi Zhan, Jeonghwan Kim, Dimitrios Bralios et al.ACL 2026 · 1 citation
- KEPo: Knowledge Evolution Poison on Graph-based Retrieval-Augmented GenerationQizhi Chen, Chao Qi, Yihong Huang, Muquan Li et al.WWW 2026
- Joint-GCG: Unified Gradient-Based Poisoning Attacks on Retrieval-Augmented Generation SystemsHaowei Wang, Rupeng Zhang, Junjie Wang, Mingyang Li et al.AAAI 2026 · 3 citations
- On the Vulnerability of Applying Retrieval-Augmented Generation within Knowledge-Intensive Application DomainsXun Xian, Ganghua Wang, Xuan Bi, Rui Zhang et al.ICML 2025
- PR-Attack: Coordinated Prompt-RAG Attacks on Retrieval-Augmented Generation in Large Language Models via Bilevel OptimizationYang Jiao, Xiaodong Wang, Kai YangSIGIR 2025 · 6 citations
