How to run a literature review in Claude Code
A worked example in Claude Code: sweep top venues, follow citations, read full text, tabulate the methods, then check every sentence against its source.
7 min read
Ask Claude Code for the related work on a topic and a tidy bibliography appears in seconds. Some of it is real. The model writes references from memory, the same way it writes everything else, so a real title can arrive with the wrong venue, a year that is off by one, or a result borrowed from a different paper. Our post on hallucinated citations covers how often that happens and why.
A better prompt does not fix this. Giving the agent a library and a method does, as long as you make it show its work. This post runs a real review from start to finish, on defenses against prompt injection, in Claude Code with Lune connected. Every table, count and quote below came out of that one session.
Connect Claude Code to the papers
Lune ships as a Claude Code plugin that bundles the MCP server with three research skills. Inside Claude Code, run:
/plugin marketplace add RetrogradeLabs/lune
/plugin install lune@retrograde-labs-lune
/mcpIn the /mcp list, choose plugin:lune:lune and sign in on the browser tab it
opens. There is no key to paste. If you would rather stay in a terminal,
npm install -g @retrograde-labs/lune-cli, then lune login and
lune install --client claude-code, connects the same tools, though without
the plugin's skills that this post uses at the end.
Claude now has twelve research tools. This review used five of them:
search_papers_many, get_paper_citations, search_papers,
extract_from_papers and verify_claims. You never have to name them. Ask in
plain English and the plugin's literature-review skill picks the right one.
Sweep before you search
One query finds one cluster of papers. Authors in neighbouring subfields describe the same idea in different words, so a review needs several angles on the same question. These are the four angles the review ran, word for word:
1. defenses that make LLM-integrated applications robust to prompt injection
attacks hidden in retrieved data or tool outputs
2. structured queries or fine-tuning that separate trusted instructions from
untrusted data to stop prompt injection
3. benchmarks and formal frameworks for evaluating prompt injection attacks
and defenses
4. prompt injection attacks against autonomous LLM agents that use tools and
browse the webAll four went out as a single search_papers_many call and came back as 15
papers from USENIX Security, IEEE S&P, CCS, NDSS, ICLR, ACL and EMNLP, published
between 2024 and 2026. Every hit says which angles found it.
Formalizing and Benchmarking Prompt Injection Attacks and Defenses
(USENIX Security 2024) ranked first for two of the four, which makes it the
natural anchor for the review. The papers that only one angle found deserve the
closer look: they show what the other queries would have missed.
Follow the citations both ways
Search finds papers that use your words. Citations find papers that build on each other whatever words they use. Software engineering calls this snowballing, and Wohlin's guidelines (EASE 2014) are the usual reference for doing it systematically: backward through a paper's references, forward through the papers that cite it.
Forward from StruQ
(USENIX Security 2025), get_paper_citations listed 46 indexed papers that
cite it. Among them were
ACE, a security architecture
for LLM-integrated apps (NDSS 2026), and
Control Illusion (AAAI 2026),
which tests whether instruction hierarchies actually hold. Neither turned up in
the sweep. One citation pass in each direction is the cheapest insurance
against writing "no prior work has done this" about a problem somebody solved
last year.
Read the papers that matter in full
Abstracts are written to sell. The numbers, the baselines and the limitations live in the body. StruQ's abstract says the system "significantly improves resistance to prompt injection attacks, with little or no impact on utility." Its evaluation gives the numbers, and adds that the defended model "is not completely immune" to the strongest attacks. That clause belongs in your related work, and the abstract never mentions it.
In this review, extract_from_papers and verify_claims did the full-text
reading. When you want to read a paper yourself, get_paper_fulltext accepts a
list of section names, so Claude can pull just the evaluation or the
limitations instead of all of it. Keep it for the handful of papers you will
quote or contrast. Each read costs a request, and skimming a full text you
never cite is a waste of both.
Put the papers side by side
The quickest way to see the shape of a literature is one table with the same
columns for every paper. extract_from_papers reads each paper's full text and
fills the columns you define. I asked for how each defense works, whether it
trains a model, the headline result as reported, and a limitation the authors
state. Here it is, condensed:
| Defense | How it works | Trains a model | Headline result, as reported |
|---|---|---|---|
| StruQ, USENIX Security 2025 | Separates prompt from data with reserved delimiters, then fine-tunes the model to follow only the prompt | Yes | TAP success on Llama falls from 97% to 9%, GCG from 97% to 58% |
| SecAlign, CCS 2025 | Preference-optimizes the model to favour the intended response over the injected one | Yes | Optimization-based attack success more than 4 times lower than under StruQ |
| DataSentinel, IEEE S&P 2025 | A detector fine-tuned as a game against adaptive attackers flags contaminated input | Yes | False positives near zero, false negatives near zero for several attacks |
| DRIP, CCS 2026 | Edits the embeddings of data tokens away from the instruction space | Yes | Attack success 66% lower than existing defenses |
| Referencing, ACL 2026 | The model tags which instruction it is executing, and answers tagged with any other are dropped | No | Attack success of 0% in some scenarios |
Five papers, five requests, and a table that took minutes rather than an afternoon. Treat it as notes and not as the final word. The last step of this review exists because a table like this one can be wrong in ways that look right.
Write by method, not paper by paper
A related-work section that walks through papers one at a time reads like an annotated bibliography. Sebastian Farquhar of Google DeepMind is blunt about it in How to Write ML Papers: "Good prior work sections are methodological." The paper-by-paper version, he adds, is fine for a first draft but should then be converted.
The table shows the methodological version at a glance. Four of the five
defenses train a model; Referencing needs no training and works through
instruction tags and output filtering. Four keep the injected instruction from
taking effect and one detects contaminated input before it reaches the model. That split is your paragraph structure, and each claim in it
already points at a paper. Advice like Farquhar's lives in Lune's guidance
library, which Claude reaches through search_research_guidance, so the agent
can apply it while it drafts.
Check every sentence before you keep it
Then I made the kind of mistakes that slip into real drafts: one number
"remembered" wrong and one generalization that sounded fair. verify_claims
checked four sentences:
- StruQ cuts TAP's success rate on Llama from 97% to 9%. Supported, with the quote "Our Llama model has significantly increased robustness against TAP (97% → 9% ASR) and GCG (97% → 58%), but is not completely immune to such attacks."
- StruQ cuts GCG's success rate on Llama below 10%. Unsupported. The same evaluation reports 97% to 58%.
- SecAlign cuts optimization-based attack success more than fourfold relative to StruQ. Supported.
- Fine-tuned defenses reliably generalize to attacks they never saw in training. Unsupported, with a verbatim quote from a CCS 2024 paper: defenses based on fine-tuning "either have limited effectiveness for prompt injection attacks that are not considered during finetuning or sacrifice generality of the LLM."
A second pass over three of the table's rows turned up something more useful than a typo. DataSentinel's near-zero false negatives hold for the attacks its authors tested. ObliInjection (NDSS 2026) ran DataSentinel and a perplexity filter against a new attack and measured high false negative rates for both. In its authors' words, they "fail to reliably detect contaminated segments crafted by ObliInjection." Both findings are true. A review that repeats only the first one misrepresents the field, and Farquhar names that as the thing that "poisons the field, annoys your colleagues, and makes reviewers angry."
When a verdict comes with a quote, the server has already checked that the quote occurs in the text it retrieved, so you can paste it into your notes without opening the PDF again.
Make it one prompt
Once the steps make sense, you can stop spelling them out. The plugin's literature-review skill encodes the same order (sweep, triage, read, check coverage), so a request like this runs the whole review:
Survey prompt injection defenses for LLM agents published since 2024.
Sweep several angles, follow citations one hop each way, tabulate mechanism,
training cost and headline result, then draft two methodological paragraphs
and verify every sentence before you show them to me.To start the skill explicitly, type /lune:lune-literature-review followed by
your topic.
What it costs, and where it stops
This review used 18 requests: four for the sweep, one for the citation pass, one to pull StruQ's abstract, five for the table and seven for the checks. That is more than the Free plan's 10 requests a day and a small slice of the 300 a day on Pro.
The corpus is deliberate about what it holds. As of September 2026 it indexes 23 computer science venues ranked by CORE and CCF, most papers in full text and most venues back to 2020. A workshop paper, a journal article or a fresh arXiv preprint sits outside it, and for those the agent should say so and search the web instead of improvising. That boundary is the point. Every paper Lune returns is one you can open, so a reference in the draft that you cannot trace back to a returned record is the one to cut.
