Introducing Spotlight: A Novel Approach for Generating Captivating Key Information from Documents
Ankan Mullick, Sombit Bose, Rounak Saha, Ayan Kumar Bhowmick, Aditya Vempaty, Prasenjit Dey, Ravi Kokku, Pawan Goyal, Niloy Ganguly
Abstract
Analyzing and processing vast amounts of textual data presents significant challenges in efficiently extracting key information. In this paper, we introduce 'Spotlight', a novel paradigm for information extraction that produces concise, engaging narratives by highlighting the most compelling aspects of a document. Unlike highlights (fragmented key points) and traditional summaries, which prioritize comprehensive coverage, spotlights selectively emphasize intriguing content to foster deeper reader engagement with the source material. We formally differentiate spotlights from related constructs and support our analysis with a detailed benchmarking study using new datasets curated for this work. To generate high-quality spotlights, we propose a twostage approach: fine-tuning a large language model on our benchmark data, followed by alignment via Direct Preference Optimization (DPO). Our comprehensive evaluation demonstrates that the resulting model not only identifies key elements with precision but also enhances readability and boosts the engagement value of the original document. Datasets and code are available at https://github.com/ ankan2/Spotlight-EMNLP2025 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b3b37f1d-bd48-4f1d-87cc-841cc509684cBuilds on16
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive SummarizationJingqing Zhang, Yao Zhao, Mohammad Saleh, Peter J. LiuICML 2020 · 2,453 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text GenerationSewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis et al.EMNLP 2023 · 225 citations
Related papers
- Better Highlighting: Creating Sub-Sentence Summary HighlightsSangwoo Cho, Kaiqiang Song, Chen Li, Dong Yu et al.EMNLP 2020 · 14 citations
- What's the Difference? Supporting Users in Identifying the Effects of Prompt and Model Changes Through Token PatternsMichael A. Hedderich, Anyi Wang, Raoyuan Zhao, Florian Eichin et al.ACL 2025
- Model-based Preference Optimization in Abstractive Summarization without Human FeedbackJaepill Choi, Kyubyung Chae, Jiwoo Song, Yohan Jo et al.EMNLP 2024 · 1 citation
- Focus, Don't Prune: Identifying Instruction-Relevant Regions for Information-Rich Image UnderstandingMincheol Kwon, Minseung Lee, Seonga Choi, Miso Choi et al.CVPR 2026 · 2 citations
- Finding Needles in Images: Can Multi-modal LLMs Locate Fine Details?Parth Thakkar, Ankush Agarwal, Prasad Kasu, Pulkit Bansal et al.ACL 2025
