PaperTrail: A Claim-Evidence Interface for Grounding Provenance in LLM-based Scholarly Q&A
Anna Martin-Boyle, Cara A. C. Leckey, Martha Brown, Harmanpreet Kaur
Abstract
Large language models (LLMs) are increasingly used in scholarly question-answering (QA) systems to help researchers synthesize vast amounts of literature. However, these systems often produce subtle errors (e.g., unsupported claims, errors of omission), and current provenance mechanisms like source citations are not granular enough for the rigorous verification that scholarly domain requires. To address this, we introduce PaperTrail, a novel interface that decomposes both LLM answers and source documents into discrete claims and evidence, mapping them to reveal supported assertions, unsupported claims, and information omitted from the source texts. We evaluated PaperTrail in a within-subjects study with 26 researchers who performed two scholarly editing tasks using PaperTrail and a baseline interface. Our results show that PaperTrail significantly lowered participants’ trust compared to the baseline. However, this increased caution did not translate to behavioral changes, as people continued to rely on LLM-generated scholarly edits to avoid a cognitively burdensome task. We discuss the value of claim-evidence matching for understanding LLM trustworthiness in scholarly settings, and present design implications for cognition-friendly communication of provenance information.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f87df7bc-9765-4eb5-b4b6-c8b08a30b74dBuilds on28
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 3,228 citations
- To Trust or to Think: Cognitive Forcing Functions Can Reduce Overreliance on AI in AI-assisted Decision-makingZana Buçinca, Maja Barbara Malaya, Krzysztof Z. GajosCSCW 2021 · 962 citations
- Questioning the AI: Informing Design Practices for Explainable AI User ExperiencesQ. Vera Liao, Daniel M. Gruen, Sarah MillerCHI 2020 · 758 citations
- Does the Whole Exceed its Parts? The Effect of AI Explanations on Complementary Team PerformanceGagan Bansal, Tongshuang Wu, Joyce Zhou, Raymond Fok et al.CHI 2021 · 713 citations
Related papers
- PaperAsk: A Benchmark for Reliability Evaluation of LLMs in Paper Search and ReadingYutao Wu, Xiao Liu, Yunhao Feng, Jiale Ding et al.WWW 2026 · 1 citation
- An Expert Schema for Evaluating Large Language Model Errors in Scholarly Question-Answering SystemsAnna Martin-Boyle, William Humphreys, Martha Brown, Cara A. C. Leckey et al.CHI 2026 · 1 citation
- LLM or Human? Perceptions of Trust and Quality in Research SummariesNil-Jana Akpinar, Sandeep Avula, Chia-Jung Lee, Brandon Dang et al.CHI 2026 · 2 citations
- Bolt-on, Verifiable Provenance for LLM-Powered Data ProcessingYiming Lin, Sepanta Zeighami, Aditya G. ParameswaranVLDB 2026
- Beyond the Chat: Executable and Verifiable Text-Editing with LLMsPhilippe Laban, Jesse Vig, Marti A. Hearst, Caiming Xiong et al.UIST 2024 · 29 citations
