Towards Faithful and Robust LLM Specialists for Evidence-Based Question-Answering
Tobias Schimanski, Jingwei Ni, Mathias Kraus, Elliott Ash, Markus Leippold
摘要
Advances towards more faithful and traceable answers of Large Language Models (LLMs) are crucial for various research and practical endeavors. One avenue in reaching this goal is basing the answers on reliable sources. However, this Evidence-Based QA has proven to work insufficiently with LLMs in terms of citing the correct sources (source quality) and truthfully representing the information within sources (answer attributability). In this work, we systematically investigate how to robustly fine-tune LLMs for better source quality and answer attributability. Specifically, we introduce a data generation pipeline with automated data quality filters, which can synthesize diversified high-quality training and testing data at scale. We further introduce four test sets to benchmark the robustness of fine-tuned specialist models. Extensive evaluation shows that fine-tuning on synthetic data improves performance on both in-and out-of-distribution. Furthermore, we show that data quality, which can be drastically improved by proposed quality filters, matters more than quantity in improving Evidence-Based QA. 1 * Equal Contributions. 1 All our codes, LLM generations, and human annotations are accessible through https://github.com/ EdisonNi-hku/Robust_Evidence_Based_QA .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Medical Graph RAG: Evidence-based Medical Large Language Model via Graph Retrieval-Augmented GenerationJunde Wu, Jiayuan Zhu, Yunli Qi, Jingkun Chen 等ACL 2025 · 被引用 64 次
- Attribution, Citation, and Quotation: A Survey of Evidence-based Text Generation with Large Language ModelsTobias Schreieder, Tim Schopf, Michael FärberACL 2026 · 被引用 10 次
- Translate Policy to Language: Flow Matching Generated Rewards for LLM ExplanationsXinyi Yang, Liang Zeng, Heng Dong, Chao Yu 等ICLR 2026 · 被引用 6 次
- HypoChainer: A Collaborative System Combining LLMs and Knowledge Graphs for Hypothesis-Driven Scientific DiscoveryHaoran Jiang, Shaohan Shi, Yunjie Yao, Chang Jiang 等IEEE VIS 2025 · 被引用 4 次
- CaseMaster: Designing and Evaluating a Probe for Oral Case Presentation Training with LLM AssistanceYang Ouyang, Yuansong Xu, Chang Jiang, Yifan Jin 等CHI 2026 · 被引用 1 次
它引用的顶会 Paper4
- FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text GenerationSewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis 等EMNLP 2023 · 被引用 225 次
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis 等EMNLP 2020 · 被引用 142 次
- Unnatural Instructions: Tuning Language Models with (Almost) No Human LaborOr Honovich, Thomas Scialom, Omer Levy, Timo SchickACL 2023 · 被引用 92 次
- Dynosaur: A Dynamic Growth Paradigm for Instruction-Tuning Data CurationDa Yin, Xiao Liu, Fan Yin, Ming Zhong 等EMNLP 2023 · 被引用 17 次
相关 Paper
- On Synthesizing Data for Context Attribution in Question AnsweringGorjan Radevski, Kiril Gashteovski, Shahbaz Syed, Christopher Malon 等ACL 2025 · 被引用 1 次
- Quality Matters: Evaluating Synthetic Data for Tool-Using LLMsShadi Iskander, Sofia Tolmach, Ori Shapira, Nachshon Cohen 等EMNLP 2024 · 被引用 2 次
- Learning from Synthetic Data Improves Multi-hop ReasoningAnmol Kabra, Yilun Yin, Albert Gong, Kamilė Stankevičiūtė 等ICLR 2026 · 被引用 6 次
- Better Datasets Start from RefineLab: Automatic Optimization for High-Quality Dataset RefinementXiaonan Luo, Yue Huang, Ping He, Xiangliang ZhangAAAI 2026 · 被引用 1 次
- Know the Known and the Unknown: Reasonable Answer Generation with Knowledge-Informed CitationsYichi Zhang, Zhuo Chen, Lingbing Guo, Jun Xu 等ACL 2026
