Attribute or Abstain: Large Language Models as Long Document Assistants
Jan Buchmann, Xiao Liu, Iryna Gurevych
Abstract
LLMs can help humans working with long documents, but are known to hallucinate. Attribution can increase trust in LLM responses: The LLM provides evidence that supports its response, which enhances verifiability. Existing approaches to attribution have only been evaluated in RAG settings, where the initial retrieval confounds LLM performance. This is crucially different from the long document setting, where retrieval is not needed, but could help. Thus, a long document specific evaluation of attribution is missing. To fill this gap, we present LAB, a benchmark of 6 diverse long document tasks with attribution, and experiments with different approaches to attribution on 5 LLMs of different sizes. We find that citation, i.e. response generation and evidence extraction in one step, performs best for large and fine-tuned models, while additional retrieval can help for small, prompted models. We investigate whether the "Lost in the Middle" phenomenon exists for attribution, but do not find this. We also find that evidence quality can predict response quality on datasets with simple responses, but not so for complex responses, as models struggle with providing evidence for complex claims. We release code and data for further investigation 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 26a9ffa5-b630-4820-8d75-512bf203c3e2Cited by top-tier papers4
- On Synthesizing Data for Context Attribution in Question AnsweringGorjan Radevski, Kiril Gashteovski, Shahbaz Syed, Christopher Malon et al.ACL 2025 · 1 citation
- Measuring and Enhancing Trustworthiness of LLMs in RAG through Grounded Attributions and Learning to RefuseMaojia Song, Shang Hong Sim, Rishabh Bhardwaj, Hai Leong Chieu et al.ICLR 2025
- In Prospect and Retrospect: Reflective Memory Management for Long-term Personalized Dialogue AgentsZhen Tan, Jun Yan, I-Hung Hsu, Rujun Han et al.ACL 2025
- LAQuer: Localized Attribution Queries in Content-grounded GenerationEran Hirsch, Aviv Slobodkin, David Wan, Elias Stengel-Eskin et al.ACL 2025
Builds on23
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive SummarizationJingqing Zhang, Yao Zhao, Mohammad Saleh, Peter J. LiuICML 2020 · 2,453 citations
Related papers
- Learning to Plan and Generate Text with CitationsConstanza Fierro, Reinald Kim Amplayo, Fantine Huot, Nicola De Cao et al.ACL 2024 · 7 citations
- Attribution, Citation, and Quotation: A Survey of Evidence-based Text Generation with Large Language ModelsTobias Schreieder, Tim Schopf, Michael FärberACL 2026 · 10 citations
- Advancing Large Language Model Attribution through Self-ImprovingLei Huang, Xiaocheng Feng, Weitao Ma, Liang Zhao et al.EMNLP 2024 · 2 citations
- Cite Pretrain: Retrieval-Free Knowledge Attribution for Large Language ModelsYukun Huang, Sanxing Chen, Jian Pei, Manzil Zaheer et al.ICLR 2026 · 1 citation
- CiteGuard: Faithful Citation Attribution for LLMs via Retrieval-Augmented ValidationYee Man Choi, Xuehang Guo, Yi R. Fung, Qingyun WangACL 2026 · 7 citations
