Learning to Plan and Generate Text with Citations
Constanza Fierro, Reinald Kim Amplayo, Fantine Huot, Nicola De Cao, Joshua Maynez, Shashi Narayan, Mirella Lapata
Abstract
The increasing demand for the deployment of LLMs in information-seeking scenarios has spurred efforts in creating verifiable systems, which generate responses to queries along with supporting evidence. In this paper, we explore the attribution capabilities of plan-based models which have been recently shown to improve the faithfulness, grounding, and controllability of generated text. We conceptualize plans as a sequence of questions which serve as blueprints of the generated content and its organization. We propose two attribution models that utilize different variants of blueprints, an abstractive model where questions are generated from scratch, and an extractive model where questions are copied from the input. Experiments on long-form question-answering show that planning consistently improves attribution quality. Moreover, the citations generated by blueprint models are more accurate compared to those obtained from LLM-based pipelines lacking a planning component.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 52fbe442-8793-463e-bd7c-2f1c2a3cfed0Cited by top-tier papers8
- CiteEval: Principle-Driven Citation Evaluation for Source AttributionYumo Xu, Peng Qi, Jifan Chen, Kunlun Liu et al.ACL 2025 · 13 citations
- CiteBART: Learning to Generate Citations for Local Citation RecommendationEge Yigit Çelik, Selma TekirEMNLP 2025 · 5 citations
- Think&Cite: Improving Attributed Text Generation with Self-Guided Tree Search and Progress Reward ModelingJunyi Li, Hwee Tou NgACL 2025 · 5 citations
- Learning to Generate Answers with Citations via Factual Consistency ModelsRami Aly, Zhiqiang Tang, Samson Tan, George KarypisACL 2024 · 2 citations
- PLANTAIN: Plan-Answer Interleaved ReasoningAnthony Liang, Jonathan Berant, Adam Fisch, Abhimanyu Goyal et al.ICML 2026 · 1 citation
Builds on8
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Adversarial NLI: A New Benchmark for Natural Language UnderstandingYixin Nie, Adina Williams, Emily Dinan, Mohit Bansal et al.ACL 2020 · 602 citations
- AmbigQA: Answering Ambiguous Open-domain QuestionsSewon Min, Julian Michael, Hannaneh Hajishirzi, Luke ZettlemoyerEMNLP 2020 · 162 citations
- Coarse-to-Fine Query Focused Multi-Document SummarizationYumo Xu, Mirella LapataEMNLP 2020 · 76 citations
- On Faithfulness and Factuality in Abstractive SummarizationJoshua Maynez, Shashi Narayan, Bernd Bohnet, Ryan T. McDonaldACL 2020 · 54 citations
Related papers
- Attribute or Abstain: Large Language Models as Long Document AssistantsJan Buchmann, Xiao Liu, Iryna GurevychEMNLP 2024
- Model Internals-based Answer Attribution for Trustworthy Retrieval-Augmented GenerationJirui Qi, Gabriele Sarti, Raquel Fernández, Arianna BisazzaEMNLP 2024 · 6 citations
- Advancing Large Language Model Attribution through Self-ImprovingLei Huang, Xiaocheng Feng, Weitao Ma, Liang Zhao et al.EMNLP 2024 · 2 citations
- Attribute First, then Generate: Locally-attributable Grounded Text GenerationAviv Slobodkin, Eran Hirsch, Arie Cattan, Tal Schuster et al.ACL 2024
- Enhancing Post-Hoc Attributions in Long Document Comprehension via Coarse Grained Answer DecompositionPritika Ramu, Koustava Goswami, Apoorv Saxena, Balaji Vasan SrinivasanEMNLP 2024 · 1 citation
