Learning to Generate Answers with Citations via Factual Consistency Models
Rami Aly, Zhiqiang Tang, Samson Tan, George Karypis
摘要
Large Language Models (LLMs) frequently hallucinate, impeding their reliability in mission-critical situations. One approach to address this issue is to provide citations to relevant sources alongside generated content, enhancing the verifiability of generations. However, citing passages accurately in answers remains a substantial challenge. This paper proposes a weakly-supervised fine-tuning method leveraging factual consistency models (FCMs). Our approach alternates between generating texts with citations and supervised fine-tuning with FCM-filtered citation data. Focused learning is integrated into the objective, directing the fine-tuning process to emphasise the factual unit tokens, as measured by an FCM. Results on the ALCE few-shot citation benchmark with various instruction-tuned LLMs demonstrate superior performance compared to in-context learning, vanilla supervised fine-tuning, and state-of-the-art methods, with an average improvement of 34.1, 15.5, and 10.5 citation F 1 points, respectively. Moreover, in a domain transfer setting we show that the obtained citation generation ability robustly transfers to unseen datasets. Notably, our citation improvements contribute to the lowest factual error rate across baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- TROVE: A Challenge for Fine-Grained Text Provenance via Source Sentence Tracing and Relationship ClassificationJunnan Zhu, Min Xiao, Yining Wang, Feifei Zhai 等ACL 2025 · 被引用 5 次
- GenProve: Learning to Generate Text with Fine-Grained ProvenanceJingxuan Wei, Xingyue Wang, Yanghaoyu Liao, Jie Dong 等ACL 2026 · 被引用 1 次
- FineRef: Fine-Grained Error Reflection and Correction for Long-Form Generation with CitationsYixing Peng, Licheng Zhang, Shancheng Fang, Yi Liu 等AAAI 2026
它引用的顶会 Paper21
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-ReflectionAkari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil 等ICLR 2024 · 被引用 1,798 次
- MAUVE: Measuring the Gap Between Neural Text and Human Text using Divergence FrontiersKrishna Pillutla, Swabha Swayamdipta, Rowan Zellers, John Thickstun 等NeurIPS 2021 · 被引用 606 次
相关 Paper
- Enabling Large Language Models to Generate Text with CitationsTianyu Gao, Howard Yen, Jiatong Yu, Danqi ChenEMNLP 2023 · 被引用 152 次
- Training Language Models to Generate Text with Citations via Fine-grained RewardsChengyu Huang, Zeqiu Wu, Yushi Hu, Wenya WangACL 2024 · 被引用 4 次
- Cite Pretrain: Retrieval-Free Knowledge Attribution for Large Language ModelsYukun Huang, Sanxing Chen, Jian Pei, Manzil Zaheer 等ICLR 2026 · 被引用 1 次
- Guidance: Sentence-Level Citation Enforcement via Prefix-Tail Guidance during LLM DecodingYirui Zhan, Xu, Jun GaoICML 2026
- Chain-of-Thought Improves Text Generation with Citations in Large Language ModelsBin Ji, Huijun Liu, Mingzhe Du, See-Kiong NgAAAI 2024 · 被引用 17 次
