Lune

NeurIPS2025Top-tier venue

NeuroGenPoisoning: Neuron-Guided Attacks on Retrieval-Augmented Generation of LLM via Genetic Optimization of External Knowledge

Hanyu Zhu, Lance Fiondella, Jiawei Yuan, Kai Zeng, Long Jiao

2025Year
8Citations
2Top-tier citations

Abstract

Retrieval-Augmented Generation (RAG) empowers Large Language Models (LLMs) to dynamically integrate external knowledge during inference, improving their factual accuracy and adaptability. However, adversaries can inject poisoned external knowledge to override the model's internal memory. While existing attacks iteratively manipulate retrieval content or prompt structure of RAG, they largely ignore the model's internal representation dynamics and neuron-level sensitivities. The underlying mechanism of RAG poisoning has not been fully studied and the effect of knowledge conflict with strong parametric knowledge in RAG is not considered. In this work, we propose NeuroGenPoisoning, a novel attack framework that generates adversarial external knowledge in RAG guided by LLM internal neuron attribution and genetic optimization. Our method first identifies a set of Poison-Responsive Neurons whose activation strongly correlates with contextual poisoning knowledge. We then employ a genetic algorithm to evolve adversarial passages that maximally activate these neurons. Crucially, our framework enables massive-scale generation of effective poisoned RAG knowledge by identifying and reusing promising but initially unsuccessful external knowledge variants via observed attribution signals. At the same time, Poison-Responsive Neurons guided poisoning can effectively resolves knowledge conflict. Experimental results across models and datasets demonstrate consistently achieving high Population Overwrite Success Rate (POSR) of over 90% while preserving fluency. Empirical evidence shows that our method effectively resolves knowledge conflict. * Corresponding Author 39th Conference on Neural Information Processing Systems (NeurIPS 2025).

override the original knowledge of LLMs, leading to hallucinations, misinformation, or targeted disinformation [7,26,27,41,50].

Prior studies, such as PoisonedRAG [54], BadRAG [45], Pandora have demonstrated that LLMs can be manipulated by injecting carefully crafted contexts in RAG. However, these attacks typically rely on pre-defined misinformation templates or manually constructed adversarial passages [2,9,13,18,45,52,54], limiting the scalability and generality of their approach. Moreover, they do not explicitly model which internal components of the LLM are responsible for context reliance or knowledge conflicts. Recent research has identified a set of context-aware neurons [32], which are responsible for integrating external content into model predictions. If such neurons can be systematically activated by poisoned knowledge, then it is possible to override a model's parametric knowledge via a targeted manipulation of its internal decision pathway.

In this work, we propose NeuroGenPoisoning, a novel attack framework that leverages Poison-Responsive Neurons, neurons that are highly sensitive to external knowledge in RAG settings. Inspired by IRCAN [32], we identify Poison-Responsive Neurons via Integrated Gradients (IG) [33], and use their activation scores as optimization signals in genetic algorithms. We begin by prompting an LLM to generate misleading external knowledge passages containing a specified incorrect answer. These adversarial seeds mimic plausible sources while embedding targeted misinformation. From this initialization, we iteratively evolve the passages using a genetic algorithm guided by neuron attribution scores, which progressively amplify their influence on the model's output. By directly optimizing for internal attribution rather than surface-level cues, NeuroGenPoisoning crafts semantically coherent and stealthy poisoned knowledge capable of overriding the LLM's internal memory. Injected into the RAG pipeline, these optimized passages consistently induce hallucinations aligned with the adversary's target, even when the model has previously memorized the correct answer. Our experiments demonstrate that NeuroGenPoisoning can efficienlty launch massive RAG poisoning under large population and consistently achieves high Population Overwrite Success Rate (POSR) in multiple open-domain question answering (QA) datasets, including SQuAD 2.0 [30], TriviaQA [19], and WikiQA [46], and a variety of LLMs such as LLaMA-2-7b [36], Vicuna-7b/13b [8], and Gemma- 7b [35]. For example, on SQuAD 2.0, our method achieves a Population Overwrite Success Rate (POSR) of over 90% on LLaMA-2-7b-chat-hf, compared to an initial POSR of about 40% to 50%. Our method proves especially effective in knowledge conflict settings, in which the model has a strong internal memory of the correct answer. We observe that as genetic optimization progresses, the distribution of query-level POSR gradually shifts to higher, indicating that more queries achieve high POSR over time. Specially, the initialized external knowledge can lead to moderate POSR, with approximately 70% of queries exhibiting POSR between 40% and 50%. However, the success remains inconsistent, with a minority of queries succeed completely (

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 3870c312-db71-46b8-9e9b-b7cf6506f501

Cited by top-tier papers2

Ask how each one uses it

Builds on24

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines