Socratic Pretraining: Question-Driven Pretraining for Controllable Summarization
Artidoro Pagnoni, Alexander R. Fabbri, Wojciech Kryscinski, Chien-Sheng Wu
Abstract
In long document controllable summarization, where labeled data is scarce, pretrained models struggle to adapt to the task and effectively respond to user queries. In this paper, we introduce Socratic pretraining, a question-driven, unsupervised pretraining objective specifically designed to improve controllability in summarization tasks. By training a model to generate and answer relevant questions in a given context, Socratic pretraining enables the model to more effectively adhere to user-provided queries and identify relevant content to be summarized. We demonstrate the effectiveness of this approach through extensive experimentation on two summarization domains, short stories and dialogue, and multiple control strategies: keywords, questions, and factoid QA pairs. Our pretraining method relies only on unlabeled documents and a question generation system and outperforms pre-finetuning approaches that use additional supervised data. Furthermore, our results show that Socratic pretraining cuts task-specific labeled data requirements in half, is more faithful to user-provided queries, and achieves state-of-the-art performance on QMSum and SQuALITY.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b901fc8e-97f5-4393-881a-8f40661eb6feCited by top-tier papers3
- Adaptive Planning for Multi-Attribute Controllable Summarization with Monte Carlo Tree SearchSangwon Ryu, Heejin Do, Yunsu Kim, Gary Geunbae Lee et al.ACL 2026 · 2 citations
- QUIDS: Query Intent Description for Exploratory Search via Dual Space ModelingYumeng Wang, Xiuying Chen, Suzan VerberneEMNLP 2025 · 1 citation
- Learning to Rank Salient Content for Query-focused SummarizationSajad Sotudeh, Nazli GoharianEMNLP 2024
Builds on11
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive SummarizationJingqing Zhang, Yao Zhao, Mohammad Saleh, Peter J. LiuICML 2020 · 2,453 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- GALAXY: A Generative Pre-trained Model for Task-Oriented Dialog with Semi-supervised Learning and Explicit Policy InjectionWanwei He, Yinpei Dai, Yinhe Zheng, Yuchuan Wu et al.AAAI 2022 · 181 citations
- Muppet: Massive Multi-task Representations with Pre-FinetuningArmen Aghajanyan, Anchit Gupta, Akshat Shrivastava, Xilun Chen et al.EMNLP 2021 · 176 citations
Related papers
- Pre-training for Abstractive Document Summarization by Reinstating Source TextYanyan Zou, Xingxing Zhang, Wei Lu, Furu Wei et al.EMNLP 2020 · 42 citations
- Masking Orchestration: Multi-Task Pretraining for Multi-Role Dialogue Representation LearningTianyi Wang, Yating Zhang, Xiaozhong Liu, Changlong Sun et al.AAAI 2020 · 8 citations
- DIONYSUS: A Pre-trained Model for Low-Resource Dialogue SummarizationYu Li, Baolin Peng, Pengcheng He, Michel Galley et al.ACL 2023 · 4 citations
- Improving Unsupervised Question Answering via Summarization-Informed Question GenerationChenyang Lyu, Lifeng Shang, Yvette Graham, Jennifer Foster et al.EMNLP 2021 · 33 citations
- KPT: Keyword-Guided Pre-training for Grounded Dialog GenerationQi Zhu, Fei Mi, Zheng Zhang, Yasheng Wang et al.AAAI 2023 · 5 citations
