GenProve: Learning to Generate Text with Fine-Grained Provenance
Jingxuan Wei, Xingyue Wang, Yanghaoyu Liao, Jie Dong, Yuchen Liu, Caijun Jia, Bihui Yu, Junnan Zhu
摘要
Large language models (LLM) often hallucinate, and while adding citations is a common solution, it is frequently insufficient for accountability as users struggle to verify how a cited source supports a generated claim. Existing methods are typically coarse-grained and fail to distinguish between direct quotes and complex reasoning. In this paper, we introduce Generation-time Fine-grained Provenance, a task where models must generate fluent answers while simultaneously producing structured, sentence-level provenance triples. To enable this, we present ReFInE (Relation-aware Fine-grained Interpretability & Evidence), a dataset featuring expert-verified annotations that distinguish between Quotation, Compression, and Inference. Building on ReFInE, we propose GenProve, a framework that combines Supervised Fine-Tuning (SFT) with Group Relative Policy Optimization (GRPO). By optimizing a composite reward for answer fidelity and provenance correctness, GenProve significantly outperforms 14 strong LLMs in joint evaluation. Crucially, our analysis uncovers a reasoning gap where models excel at surfacelevel quotation but struggle significantly with inference-based provenance, suggesting that verifiable reasoning remains a frontier challenge distinct from surface-level citation. The code is publicly available at https://github. com/yanghaoyuliao/GenProve .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- Enabling Large Language Models to Generate Text with CitationsTianyu Gao, Howard Yen, Jiatong Yu, Danqi ChenEMNLP 2023 · 被引用 152 次
- Automatic Generation of Citation Texts in Scholarly Papers: A Pilot StudyXinyu Xing, Xiaosheng Fan, Xiaojun WanACL 2020 · 被引用 39 次
- Achieving Human Parity in Content-Grounded Datasets GenerationAsaf Yehudai, Boaz Carmeli, Yosi Mass, Ofir Arviv 等ICLR 2024 · 被引用 9 次
- MIRAGE: Scaling Test-Time Inference with Parallel Graph-Retrieval-Augmented Reasoning ChainsKaiwen Wei, Rui Shan, Dongsheng Zou, Jianzhong Yang 等AAAI 2026 · 被引用 5 次
相关 Paper
- TROVE: A Challenge for Fine-Grained Text Provenance via Source Sentence Tracing and Relationship ClassificationJunnan Zhu, Min Xiao, Yining Wang, Feifei Zhai 等ACL 2025 · 被引用 5 次
- Training Language Models to Generate Text with Citations via Fine-grained RewardsChengyu Huang, Zeqiu Wu, Yushi Hu, Wenya WangACL 2024 · 被引用 4 次
- Towards Verifiable Text Generation with Evolving Memory and Self-ReflectionHao Sun, Hengyi Cai, Bo Wang, Yingyan Hou 等EMNLP 2024 · 被引用 6 次
- DecepChain: Inducing Deceptive Reasoning in Large Language ModelsWei Shen, Han Wang, Haoyu Li, Huan ZhangICML 2026 · 被引用 4 次
- Learning to Reason for Hallucination Span DetectionHsuan Su, Ting-Yao Hu, Hema Swetha Koppula, Kundan Krishna 等ICLR 2026 · 被引用 8 次
