Expository Text Generation: Imitate, Retrieve, Paraphrase
Nishant Balepur, Jie Huang, Kevin Chen-Chuan Chang
Abstract
Expository documents are vital resources for conveying complex information to readers. Despite their usefulness, writing expository text by hand is a challenging process that requires careful content planning, obtaining facts from multiple sources, and the ability to clearly synthesize these facts. To ease these burdens, we propose the task of expository text generation, which seeks to automatically generate an accurate and stylistically consistent expository text for a topic by intelligently searching a knowledge source. We solve our task by developing IRP, a framework that overcomes the limitations of retrieval-augmented models and iteratively performs content planning, fact retrieval, and rephrasing. Through experiments on three diverse, newly-collected datasets, we show that IRP produces factual and organized expository texts that accurately inform readers. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f1205d2a-7a8e-4cb1-bb5f-ff69ad00546cCited by top-tier papers11
- AutoSurvey: Large Language Models Can Automatically Write SurveysYidong Wang, Qi Guo, Wenjin Yao, Hongbo Zhang et al.NeurIPS 2024 · 151 citations
- Into the Unknown Unknowns: Engaged Human Learning through Participation in Language Model Agent ConversationsYucheng Jiang, Yijia Shao, Dekun Ma, Sina J. Semnani et al.EMNLP 2024 · 8 citations
- WikiAutoGen: Towards Multi-Modal Wikipedia-Style Article GenerationZhongyu Yang, Jun Chen, Dannong Xu, Junjie Fei et al.ICCV 2025 · 3 citations
- Writing Like the Best: Exemplar-Based Expository Text GenerationYuxiang Liu, Kevin Chen-Chuan ChangACL 2025 · 2 citations
- Language Models Don't Know What You Want: Evaluating Personalization in Deep Research Needs Real UsersNishant Balepur, Malachi Hamada, Varsha Kishore, Sergey Feldman et al.ACL 2026 · 1 citation
Builds on12
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- Active Retrieval Augmented GenerationZhengbao Jiang, Frank F. Xu, Luyu Gao, Zhiqing Sun et al.EMNLP 2023 · 315 citations
- Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step QuestionsHarsh Trivedi, Niranjan Balasubramanian, Tushar Khot, Ashish SabharwalACL 2023 · 187 citations
- RocketQAv2: A Joint Training Method for Dense Passage Retrieval and Passage Re-rankingRuiyang Ren, Yingqi Qu, Jing Liu, Wayne Xin Zhao et al.EMNLP 2021 · 147 citations
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis et al.EMNLP 2020 · 142 citations
Related papers
- Analysis of Plan-based Retrieval for Grounded Text GenerationAmeya Godbole, Nicholas Monath, Seungyeon Kim, Ankit Singh Rawat et al.EMNLP 2024
- CiteBench: A Benchmark for Scientific Citation Text GenerationMartin Funkquist, Ilia Kuznetsov, Yufang Hou, Iryna GurevychEMNLP 2023 · 2 citations
- EventRAG: Enhancing LLM Generation with Event Knowledge GraphsZairun Yang, Yilin Wang, Zhengyan Shi, Yuan Yao et al.ACL 2025 · 6 citations
- Paper2Rebuttal: A Multi-Agent Framework for Transparent Author Response AssistanceQianli Ma, Chang Guo, Zhiheng Tian, Siyu Wang et al.ACL 2026 · 6 citations
- ReviewRL: Towards Automated Scientific Review with RLSihang Zeng, Kai Tian, Kaiyan Zhang, Yuru Wang et al.EMNLP 2025
