PaperTrail: A Claim-Evidence Interface for Grounding Provenance in LLM-based Scholarly Q&A
Anna Martin-Boyle, Cara A. C. Leckey, Martha Brown, Harmanpreet Kaur
摘要
Large language models (LLMs) are increasingly used in scholarly question-answering (QA) systems to help researchers synthesize vast amounts of literature. However, these systems often produce subtle errors (e.g., unsupported claims, errors of omission), and current provenance mechanisms like source citations are not granular enough for the rigorous verification that scholarly domain requires. To address this, we introduce PaperTrail, a novel interface that decomposes both LLM answers and source documents into discrete claims and evidence, mapping them to reveal supported assertions, unsupported claims, and information omitted from the source texts. We evaluated PaperTrail in a within-subjects study with 26 researchers who performed two scholarly editing tasks using PaperTrail and a baseline interface. Our results show that PaperTrail significantly lowered participants’ trust compared to the baseline. However, this increased caution did not translate to behavioral changes, as people continued to rely on LLM-generated scholarly edits to avoid a cognitively burdensome task. We discuss the value of claim-evidence matching for understanding LLM trustworthiness in scholarly settings, and present design implications for cognition-friendly communication of provenance information.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper28
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 被引用 3,228 次
- To Trust or to Think: Cognitive Forcing Functions Can Reduce Overreliance on AI in AI-assisted Decision-makingZana Buçinca, Maja Barbara Malaya, Krzysztof Z. GajosCSCW 2021 · 被引用 962 次
- Questioning the AI: Informing Design Practices for Explainable AI User ExperiencesQ. Vera Liao, Daniel M. Gruen, Sarah MillerCHI 2020 · 被引用 758 次
- Does the Whole Exceed its Parts? The Effect of AI Explanations on Complementary Team PerformanceGagan Bansal, Tongshuang Wu, Joyce Zhou, Raymond Fok 等CHI 2021 · 被引用 713 次
相关 Paper
- PaperAsk: A Benchmark for Reliability Evaluation of LLMs in Paper Search and ReadingYutao Wu, Xiao Liu, Yunhao Feng, Jiale Ding 等WWW 2026 · 被引用 1 次
- An Expert Schema for Evaluating Large Language Model Errors in Scholarly Question-Answering SystemsAnna Martin-Boyle, William Humphreys, Martha Brown, Cara A. C. Leckey 等CHI 2026 · 被引用 1 次
- LLM or Human? Perceptions of Trust and Quality in Research SummariesNil-Jana Akpinar, Sandeep Avula, Chia-Jung Lee, Brandon Dang 等CHI 2026 · 被引用 2 次
- Bolt-on, Verifiable Provenance for LLM-Powered Data ProcessingYiming Lin, Sepanta Zeighami, Aditya G. ParameswaranVLDB 2026
- Beyond the Chat: Executable and Verifiable Text-Editing with LLMsPhilippe Laban, Jesse Vig, Marti A. Hearst, Caiming Xiong 等UIST 2024 · 被引用 29 次
