# How to run a literature review in Claude Code

A worked example in Claude Code: sweep top venues, follow citations, read full text, tabulate the methods, then check every sentence against its source.

Published 2026-10-01

Ask Claude Code for the related work on a topic and a tidy bibliography appears
in seconds. Some of it is real. The model writes references from memory, the
same way it writes everything else, so a real title can arrive with the wrong
venue, a year that is off by one, or a result borrowed from a different paper.
[Our post on hallucinated citations](https://luneresearch.com/blog/ai-hallucinated-citations) covers how
often that happens and why.

A better prompt does not fix this. Giving the agent a library and a method
does, as long as you make it show its work. This post runs a real review from
start to finish, on defenses against prompt injection, in Claude Code with Lune
connected. Every table, count and quote below came out of that one session.

## Connect Claude Code to the papers

Lune ships as a Claude Code plugin that bundles the MCP server with three
research skills. Inside Claude Code, run:

```
/plugin marketplace add RetrogradeLabs/lune
/plugin install lune@retrograde-labs-lune
/mcp
```

In the `/mcp` list, choose `plugin:lune:lune` and sign in on the browser tab it
opens. There is no key to paste. If you would rather stay in a terminal,
`npm install -g @retrograde-labs/lune-cli`, then `lune login` and
`lune install --client claude-code`, connects the same tools, though without
the plugin's skills that this post uses at the end.

Claude now has twelve research tools. This review used five of them:
`search_papers_many`, `get_paper_citations`, `search_papers`,
`extract_from_papers` and `verify_claims`. You never have to name them. Ask in
plain English and the plugin's literature-review skill picks the right one.

## Sweep before you search

One query finds one cluster of papers. Authors in neighbouring subfields
describe the same idea in different words, so a review needs several angles on
the same question. These are the four angles the review ran, word for word:

```
1. defenses that make LLM-integrated applications robust to prompt injection
   attacks hidden in retrieved data or tool outputs
2. structured queries or fine-tuning that separate trusted instructions from
   untrusted data to stop prompt injection
3. benchmarks and formal frameworks for evaluating prompt injection attacks
   and defenses
4. prompt injection attacks against autonomous LLM agents that use tools and
   browse the web
```

All four went out as a single `search_papers_many` call and came back as 15
papers from USENIX Security, IEEE S&P, CCS, NDSS, ICLR, ACL and EMNLP, published
between 2024 and 2026. Every hit says which angles found it.
[Formalizing and Benchmarking Prompt Injection Attacks and Defenses](https://luneresearch.com/papers/c9787b94-61a0-4e72-9dc7-ace327ebac77)
(USENIX Security 2024) ranked first for two of the four, which makes it the
natural anchor for the review. The papers that only one angle found deserve the
closer look: they show what the other queries would have missed.

## Follow the citations both ways

Search finds papers that use your words. Citations find papers that build on
each other whatever words they use. Software engineering calls this
snowballing, and
[Wohlin's guidelines](https://doi.org/10.1145/2601248.2601268) (EASE 2014) are
the usual reference for doing it systematically: backward through a paper's
references, forward through the papers that cite it.

Forward from [StruQ](https://luneresearch.com/papers/549eb273-9f08-43b7-afc7-bba2111b870b)
(USENIX Security 2025), `get_paper_citations` listed 46 indexed papers that
cite it. Among them were
[ACE](https://luneresearch.com/papers/961d2212-4de1-477c-8576-6a3ce730d4f4), a security architecture
for LLM-integrated apps (NDSS 2026), and
[Control Illusion](https://luneresearch.com/papers/434464ef-d75b-4878-b6a7-29d475ae1d1a) (AAAI 2026),
which tests whether instruction hierarchies actually hold. Neither turned up in
the sweep. One citation pass in each direction is the cheapest insurance
against writing "no prior work has done this" about a problem somebody solved
last year.

## Read the papers that matter in full

Abstracts are written to sell. The numbers, the baselines and the limitations
live in the body. StruQ's abstract says the system "significantly improves
resistance to prompt injection attacks, with little or no impact on utility."
Its evaluation gives the numbers, and adds that the defended model "is not
completely immune" to the strongest attacks. That clause belongs in your
related work, and the abstract never mentions it.

In this review, `extract_from_papers` and `verify_claims` did the full-text
reading. When you want to read a paper yourself, `get_paper_fulltext` accepts a
list of section names, so Claude can pull just the evaluation or the
limitations instead of all of it. Keep it for the handful of papers you will
quote or contrast. Each read costs a request, and skimming a full text you
never cite is a waste of both.

## Put the papers side by side

The quickest way to see the shape of a literature is one table with the same
columns for every paper. `extract_from_papers` reads each paper's full text and
fills the columns you define. I asked for how each defense works, whether it
trains a model, the headline result as reported, and a limitation the authors
state. Here it is, condensed:

| Defense                                                                     | How it works                                                                                             | Trains a model | Headline result, as reported                                               |
| --------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------- | -------------- | -------------------------------------------------------------------------- |
| [StruQ](https://luneresearch.com/papers/549eb273-9f08-43b7-afc7-bba2111b870b), USENIX Security 2025 | Separates prompt from data with reserved delimiters, then fine-tunes the model to follow only the prompt | Yes            | TAP success on Llama falls from 97% to 9%, GCG from 97% to 58%             |
| [SecAlign](https://luneresearch.com/papers/1eb8f454-16fc-470b-9b13-3f0e4f58be3c), CCS 2025          | Preference-optimizes the model to favour the intended response over the injected one                     | Yes            | Optimization-based attack success more than 4 times lower than under StruQ |
| [DataSentinel](https://luneresearch.com/papers/c42c511e-b7be-4d96-b9af-1ec253382938), IEEE S&P 2025 | A detector fine-tuned as a game against adaptive attackers flags contaminated input                      | Yes            | False positives near zero, false negatives near zero for several attacks   |
| [DRIP](https://luneresearch.com/papers/830316da-3fe9-4d80-a4f9-954316ac57f1), CCS 2026              | Edits the embeddings of data tokens away from the instruction space                                      | Yes            | Attack success 66% lower than existing defenses                            |
| [Referencing](https://luneresearch.com/papers/3129dcb0-838c-4737-b410-42e7b166af14), ACL 2026       | The model tags which instruction it is executing, and answers tagged with any other are dropped          | No             | Attack success of 0% in some scenarios                                     |

Five papers, five requests, and a table that took minutes rather than an
afternoon. Treat it as notes and not as the final word. The last step of this
review exists because a table like this one can be wrong in ways that look
right.

## Write by method, not paper by paper

A related-work section that walks through papers one at a time reads like an
annotated bibliography. Sebastian Farquhar of Google DeepMind is blunt about it
in
[How to Write ML Papers](https://sebastianfarquhar.com/on-research/2024/11/04/how_to_write_ml_papers/):
"Good prior work sections are methodological." The paper-by-paper version, he
adds, is fine for a first draft but should then be converted.

The table shows the methodological version at a glance. Four of the five
defenses train a model; Referencing needs no training and works through
instruction tags and output filtering. Four keep the injected instruction from
taking effect and one detects contaminated input before it reaches the model. That split is your paragraph structure, and each claim in it
already points at a paper. Advice like Farquhar's lives in Lune's guidance
library, which Claude reaches through `search_research_guidance`, so the agent
can apply it while it drafts.

## Check every sentence before you keep it

Then I made the kind of mistakes that slip into real drafts: one number
"remembered" wrong and one generalization that sounded fair. `verify_claims`
checked four sentences:

- **StruQ cuts TAP's success rate on Llama from 97% to 9%.** Supported, with
  the quote "Our Llama model has significantly increased robustness against TAP
  (97% → 9% ASR) and GCG (97% → 58%), but is not completely immune to such
  attacks."
- **StruQ cuts GCG's success rate on Llama below 10%.** Unsupported. The same
  evaluation reports 97% to 58%.
- **SecAlign cuts optimization-based attack success more than fourfold
  relative to StruQ.** Supported.
- **Fine-tuned defenses reliably generalize to attacks they never saw in
  training.** Unsupported, with a verbatim quote from
  [a CCS 2024 paper](https://luneresearch.com/papers/d029a7fd-c200-43b5-af89-e205530e9986): defenses
  based on fine-tuning "either have limited effectiveness for prompt injection
  attacks that are not considered during finetuning or sacrifice generality of
  the LLM."

A second pass over three of the table's rows turned up something more useful
than a typo.
DataSentinel's near-zero false negatives hold for the attacks its authors
tested. [ObliInjection](https://luneresearch.com/papers/e180c011-6c8f-4fa4-bfb7-f77501530277) (NDSS 2026)
ran DataSentinel and a perplexity filter against a new attack and measured high
false negative rates for both. In its authors' words, they "fail to reliably
detect contaminated segments crafted by ObliInjection." Both findings are true.
A review that repeats only the first one misrepresents the field, and Farquhar
names that as the thing that "poisons the field, annoys your colleagues, and
makes reviewers angry."

When a verdict comes with a quote, the server has already checked that the
quote occurs in the text it retrieved, so you can paste it into your notes
without opening the PDF again.

## Make it one prompt

Once the steps make sense, you can stop spelling them out. The plugin's
literature-review skill encodes the same order (sweep, triage, read, check
coverage), so a request like this runs the whole review:

```
Survey prompt injection defenses for LLM agents published since 2024.
Sweep several angles, follow citations one hop each way, tabulate mechanism,
training cost and headline result, then draft two methodological paragraphs
and verify every sentence before you show them to me.
```

To start the skill explicitly, type `/lune:lune-literature-review` followed by
your topic.

## What it costs, and where it stops

This review used 18 requests: four for the sweep, one for the citation pass, one
to pull StruQ's abstract, five for the table and seven for the checks. That is
more than the Free plan's 10 requests a
day and a small slice of the 300 a day on
Pro.

The corpus is deliberate about what it holds. As of September 2026 it indexes
23 computer science venues ranked by CORE and CCF, most papers in full text and
most venues back to 2020. A workshop paper, a journal article or a fresh arXiv
preprint sits outside it, and for those the agent should say so and search the
web instead of improvising. That boundary is the point. Every paper Lune
returns is one you can open, so a reference in the draft that you cannot trace
back to a returned record is the one to cut.
