A Reality Check on Context Utilisation for Retrieval-Augmented Generation
Lovisa Hagström, Sara Vera Marjanovic, Haeun Yu, Arnav Arora, Christina Lioma, Maria Maistro, Pepa Atanasova, Isabelle Augenstein
Abstract
Retrieval-augmented generation (RAG) helps address the limitations of parametric knowledge embedded within a language model (LM). In real world settings, retrieved information can vary in complexity, yet most investigations of LM utilisation of context has been limited to synthetic text. We introduce DRUID (Dataset of Retrieved Unreliable, Insufficient and Difficultto-understand contexts) with real-world queries and contexts manually annotated for stance. The dataset is based on the prototypical task of automated claim verification, for which automated retrieval of real-world evidence is crucial. We compare DRUID to synthetic datasets (CounterFact, ConflictQA) and find that artificial datasets often fail to represent the complexity and diversity of realistically retrieved context. We show that synthetic datasets exaggerate context characteristics rare in real retrieved data, which leads to inflated context utilisation results, as measured by our novel ACU score. Moreover, while previous work has mainly focused on singleton context characteristics to explain context utilisation, correlations between singleton context properties and ACU on DRUID are surprisingly small compared to other properties related to context source. Overall, our work underscores the need for realworld aligned context utilisation studies to represent and improve performance in real-world RAG settings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9eb279d7-c336-45c5-b18d-8573a501408dCited by top-tier papers4
- Show Me the Work: Fact-Checkers' Requirements for Explainable Automated Fact-CheckingGreta Warren, Irina Shklovski, Isabelle AugensteinCHI 2025 · 15 citations
- Accommodate Knowledge Conflicts in Retrieval-augmented LLMs: Towards Robust Response Generation in the WildJiatai Wang, Zhiwei Xu, Di Jin, Xuewen Yang et al.AAAI 2026
- How to Compare Things Properly? A Study of Argument Relevance in Comparative Question AnsweringIrina Nikishina, Saba Anwar, Nikolay Dolgov, Maria Manina et al.ACL 2025
- How Context Shapes Truth: Geometric Transformations of Statement-level Truth Representations in LLMsShivam Adarsh, Maria Maistro, Christina LiomaACL 2026
Builds on25
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 3,415 citations
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
- Large Language Models Can Be Easily Distracted by Irrelevant ContextFreda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales et al.ICML 2023 · 970 citations
- Adaptive Chameleon or Stubborn Sloth: Revealing the Behavior of Large Language Models in Knowledge ConflictsJian Xie, Kai Zhang, Jiangjie Chen, Renze Lou et al.ICLR 2024 · 294 citations
- Dense X Retrieval: What Retrieval Granularity Should We Use?Tong Chen, Hongwei Wang, Sihao Chen, Wenhao Yu et al.EMNLP 2024 · 52 citations
Related papers
- CUB: Benchmarking Context Utilisation Techniques for Language ModelsLovisa Hagström, Youna Kim, Haeun Yu, Sang-goo Lee et al.ACL 2026 · 5 citations
- Conflict-Aware Soft Prompting for Retrieval-Augmented GenerationEunseong Choi, June Park, Hyeri Lee, Jongwuk LeeEMNLP 2025 · 1 citation
- LLMs Trust Humans More, That's a Problem! Unveiling and Mitigating the Authority Bias in Retrieval-Augmented GenerationYuxuan Li, Xinwei Guo, Jiashi Gao, Guanhua Chen et al.ACL 2025
- Medical Graph RAG: Evidence-based Medical Large Language Model via Graph Retrieval-Augmented GenerationJunde Wu, Jiayuan Zhu, Yunli Qi, Jingkun Chen et al.ACL 2025 · 64 citations
- Astute RAG: Overcoming Imperfect Retrieval Augmentation and Knowledge Conflicts for Large Language ModelsFei Wang, Xingchen Wan, Ruoxi Sun, Jiefeng Chen et al.ACL 2025 · 50 citations
