PDSum: Prototype-driven Continuous Summarization of Evolving Multi-document Sets Stream
Susik Yoon, Hou Pong Chan, Jiawei Han
摘要
Summarizing text-rich documents has been long studied in the literature, but most of the existing efforts have been made to summarize a static and predefined multi-document set. With the rapid development of online platforms for generating and distributing text-rich documents, there arises an urgent need for continuously summarizing dynamically evolving multi-document sets where the composition of documents and sets is changing over time. This is especially challenging as the summarization should be not only effective in incorporating relevant, novel, and distinctive information from each concurrent multi-document set, but also efficient in serving online applications. In this work, we propose a new summarization problem, Evolving Multi-Document sets stream Summarization (EMDS), and introduce a novel unsupervised algorithm PDSum with the idea of prototype-driven continuous summarization. PDSum builds a lightweight prototype of each multi-document set and exploits it to adapt to new documents while preserving accumulated knowledge from previous documents. To update new summaries, the most representative sentences for each multi-document set are extracted by measuring their similarities to the prototypes. A thorough evaluation with real multi-document sets streams demonstrates that PDSum outperforms state-of-the-art unsupervised multi-document summarization algorithms in EMDS in terms of relevance, novelty, and distinctiveness and is also robust to various evaluation settings.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Unsupervised Story Discovery from Continuous News Streams via Scalable Thematic EmbeddingSusik Yoon, Dongha Lee, Yunyi Zhang, Jiawei HanSIGIR 2023 · 被引用 8 次
- Online Drift Detection with Maximum Concept DiscrepancyKe Wan, Yi Liang, Susik YoonKDD 2024 · 被引用 4 次
- Can LMs Generalize to Future Data? An Empirical Analysis on Text SummarizationChi Seng Cheang, Hou Pong Chan, Derek F. Wong, Xuebo Liu 等EMNLP 2023 · 被引用 2 次
- CREAM: Continual Retrieval on Dynamic Streaming Corpora with Adaptive Soft MemoryHuiJeong Son, Hyeongu Kang, Sunho Kim, Subeen Ho 等KDD 2026 · 被引用 1 次
- Can Structural Cues Save LLMs? Evaluating Language Models in Massive Document StreamsYukyung Lee, Yebin Lim, Woojun Jung, Wonjun Choi 等KDD 2026
它引用的顶会 Paper20
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 被引用 2,496 次
- PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive SummarizationJingqing Zhang, Yao Zhao, Mohammad Saleh, Peter J. LiuICML 2020 · 被引用 2,453 次
- Extractive Summarization as Text MatchingMing Zhong, Pengfei Liu, Yiran Chen, Danqing Wang 等ACL 2020 · 被引用 410 次
- Heterogeneous Graph Neural Networks for Extractive Document SummarizationDanqing Wang, Pengfei Liu, Yining Zheng, Xipeng Qiu 等ACL 2020 · 被引用 275 次
相关 Paper
- Be Relevant, Non-Redundant, and Timely: Deep Reinforcement Learning for Real-Time Event SummarizationMin Yang, Chengming Li, Fei Sun, Zhou Zhao 等AAAI 2020 · 被引用 8 次
- Compressed Heterogeneous Graph for Abstractive Multi-Document SummarizationMiao Li, Jianzhong Qi, Jey Han LauAAAI 2023 · 被引用 14 次
- PRIMERA: Pyramid-based Masked Sentence Pre-training for Multi-document SummarizationWen Xiao, Iz Beltagy, Giuseppe Carenini, Arman CohanACL 2022 · 被引用 147 次
- SgSum: Transforming Multi-document Summarization into Sub-graph SelectionMoye Chen, Wei Li, Jiachen Liu, Xinyan Xiao 等EMNLP 2021 · 被引用 21 次
- HRSTORY: Historical News Review Based Online Story DiscoveryRenjie Zhou, Haoran Ye, Jian Wan, Yong LiaoKDD 2025
