Towards Faithful and Robust LLM Specialists for Evidence-Based Question-Answering
Tobias Schimanski, Jingwei Ni, Mathias Kraus, Elliott Ash, Markus Leippold
Abstract
Advances towards more faithful and traceable answers of Large Language Models (LLMs) are crucial for various research and practical endeavors. One avenue in reaching this goal is basing the answers on reliable sources. However, this Evidence-Based QA has proven to work insufficiently with LLMs in terms of citing the correct sources (source quality) and truthfully representing the information within sources (answer attributability). In this work, we systematically investigate how to robustly fine-tune LLMs for better source quality and answer attributability. Specifically, we introduce a data generation pipeline with automated data quality filters, which can synthesize diversified high-quality training and testing data at scale. We further introduce four test sets to benchmark the robustness of fine-tuned specialist models. Extensive evaluation shows that fine-tuning on synthetic data improves performance on both in-and out-of-distribution. Furthermore, we show that data quality, which can be drastically improved by proposed quality filters, matters more than quantity in improving Evidence-Based QA. 1 * Equal Contributions. 1 All our codes, LLM generations, and human annotations are accessible through https://github.com/ EdisonNi-hku/Robust_Evidence_Based_QA .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 15deadfe-2867-4fc0-9fd8-b88a8692a309Cited by top-tier papers9
- Medical Graph RAG: Evidence-based Medical Large Language Model via Graph Retrieval-Augmented GenerationJunde Wu, Jiayuan Zhu, Yunli Qi, Jingkun Chen et al.ACL 2025 · 64 citations
- Attribution, Citation, and Quotation: A Survey of Evidence-based Text Generation with Large Language ModelsTobias Schreieder, Tim Schopf, Michael FärberACL 2026 · 10 citations
- Translate Policy to Language: Flow Matching Generated Rewards for LLM ExplanationsXinyi Yang, Liang Zeng, Heng Dong, Chao Yu et al.ICLR 2026 · 6 citations
- HypoChainer: A Collaborative System Combining LLMs and Knowledge Graphs for Hypothesis-Driven Scientific DiscoveryHaoran Jiang, Shaohan Shi, Yunjie Yao, Chang Jiang et al.IEEE VIS 2025 · 4 citations
- CaseMaster: Designing and Evaluating a Probe for Oral Case Presentation Training with LLM AssistanceYang Ouyang, Yuansong Xu, Chang Jiang, Yifan Jin et al.CHI 2026 · 1 citation
Builds on4
- FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text GenerationSewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis et al.EMNLP 2023 · 225 citations
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis et al.EMNLP 2020 · 142 citations
- Unnatural Instructions: Tuning Language Models with (Almost) No Human LaborOr Honovich, Thomas Scialom, Omer Levy, Timo SchickACL 2023 · 92 citations
- Dynosaur: A Dynamic Growth Paradigm for Instruction-Tuning Data CurationDa Yin, Xiao Liu, Fan Yin, Ming Zhong et al.EMNLP 2023 · 17 citations
Related papers
- On Synthesizing Data for Context Attribution in Question AnsweringGorjan Radevski, Kiril Gashteovski, Shahbaz Syed, Christopher Malon et al.ACL 2025 · 1 citation
- Quality Matters: Evaluating Synthetic Data for Tool-Using LLMsShadi Iskander, Sofia Tolmach, Ori Shapira, Nachshon Cohen et al.EMNLP 2024 · 2 citations
- Learning from Synthetic Data Improves Multi-hop ReasoningAnmol Kabra, Yilun Yin, Albert Gong, Kamilė Stankevičiūtė et al.ICLR 2026 · 6 citations
- Better Datasets Start from RefineLab: Automatic Optimization for High-Quality Dataset RefinementXiaonan Luo, Yue Huang, Ping He, Xiangliang ZhangAAAI 2026 · 1 citation
- Know the Known and the Unknown: Reasonable Answer Generation with Knowledge-Informed CitationsYichi Zhang, Zhuo Chen, Lingbing Guo, Jun Xu et al.ACL 2026
