On Synthesizing Data for Context Attribution in Question Answering
Gorjan Radevski, Kiril Gashteovski, Shahbaz Syed, Christopher Malon, Sebastien Nicolas, Chia-Chien Hung, Timo Sztyler, Verena Heußer, Wiem Ben Rim, Masafumi Enomoto, Kunihiro Takeoka, Masafumi Oyamada
Abstract
Question Answering (QA) accounts for a significant portion of LLM usage "in the wild". However, LLMs sometimes produce false or misleading responses, also known as hallucinations. Therefore, grounding the generated answers in contextually provided informationi.e., providing evidence for the generated textis paramount for LLMs' trustworthiness. Providing this information is the task of context attribution. In this paper, we systematically study LLM-based approaches for this task, namely we investigate (i) zero-shot inference, (ii) LLM ensembling, and (iii) fine-tuning of small LMs on synthetic data generated by larger LLMs. Our key contribution is SYNQA: a novel generative strategy for synthesizing context attribution data. Given selected context sentences, an LLM generates QA pairs that are supported by these sentences. This leverages LLMs' natural strengths in text generation while ensuring clear attribution paths in the synthetic training data. We show that the attribution data synthesized via SYNQA is highly effective for fine-tuning small LMs for context attribution in different QA tasks and domains. Finally, with a user study, we validate the usefulness of small, efficient LMs (fine-tuned on synthetic data from SYNQA) in context attribution for QA.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 48243dbb-fe9b-4a4b-a36f-d1f7e49c3983Builds on18
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- On the Exploitability of Instruction TuningManli Shu, Jiongxiao Wang, Chen Zhu, Jonas Geiping et al.NeurIPS 2023 · 166 citations
- Hierarchical Graph Network for Multi-hop Question AnsweringYuwei Fang, Siqi Sun, Zhe Gan, Rohit Pillai et al.EMNLP 2020 · 157 citations
- Select, Answer and Explain: Interpretable Multi-Hop Reading Comprehension over Multiple DocumentsMing Tu, Kevin Huang, Guangtao Wang, Jing Huang et al.AAAI 2020 · 155 citations
- Large Dual Encoders Are Generalizable RetrieversJianmo Ni, Chen Qu, Jing Lu, Zhuyun Dai et al.EMNLP 2022 · 145 citations
Related papers
- Towards Faithful and Robust LLM Specialists for Evidence-Based Question-AnsweringTobias Schimanski, Jingwei Ni, Mathias Kraus, Elliott Ash et al.ACL 2024
- Attribute or Abstain: Large Language Models as Long Document AssistantsJan Buchmann, Xiao Liu, Iryna GurevychEMNLP 2024
- Advancing Large Language Model Attribution through Self-ImprovingLei Huang, Xiaocheng Feng, Weitao Ma, Liang Zhao et al.EMNLP 2024 · 2 citations
- Teaching Language Models to Hallucinate Less with Synthetic TasksErik Jones, Hamid Palangi, Clarisse Simões, Varun Chandrasekaran et al.ICLR 2024 · 43 citations
- Multi-Level Explanations for Generative Language ModelsLucas Monteiro Paes, Dennis Wei, Hyo Jin Do, Hendrik Strobelt et al.ACL 2025 · 16 citations
