SCStory: Self-supervised and Continual Online Story Discovery
Susik Yoon, Yu Meng, Dongha Lee, Jiawei Han
Abstract
We present a framework SCStory for online story discovery, that helps people digest rapidly published news article streams in real-time without human annotations. To organize news article streams into stories, existing approaches directly encode the articles and cluster them based on representation similarity. However, these methods yield noisy and inaccurate story discovery results because the generic article embeddings do not effectively reflect the story-indicative semantics in an article and cannot adapt to the rapidly evolving news article streams. SCStory employs self-supervised and continual learning with a novel idea of story-indicative adaptive modeling of news article streams. With a lightweight hierarchical embedding module that first learns sentence representations and then article representations, SCStory identifies story-relevant information of news articles and uses them to discover stories. The embedding module is continuously updated to adapt to evolving news streams with a contrastive learning objective, backed up by two unique techniques, confidence-aware memory replay and prioritized-augmentation, employed for label absence and data scarcity problems. Thorough experiments on real and the latest news data sets demonstrate that SCStory outperforms existing state-of-the-art algorithms for unsupervised online story discovery.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0b50067f-2ced-44ea-86ca-25c0b6e5683eCited by top-tier papers4
- Online Drift Detection with Maximum Concept DiscrepancyKe Wan, Yi Liang, Susik YoonKDD 2024 · 4 citations
- CREAM: Continual Retrieval on Dynamic Streaming Corpora with Adaptive Soft MemoryHuiJeong Son, Hyeongu Kang, Sunho Kim, Subeen Ho et al.KDD 2026 · 1 citation
- Can Structural Cues Save LLMs? Evaluating Language Models in Massive Document StreamsYukyung Lee, Yebin Lim, Woojun Jung, Wonjun Choi et al.KDD 2026
- Analyzing Temporal Complex Events with Large Language Models? A Benchmark towards Temporal, Long Context UnderstandingZhihan Zhang, Yixin Cao, Chenchen Ye, Yunshan Ma et al.ACL 2024
Builds on13
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 2,496 citations
- Understanding Contrastive Representation Learning through Alignment and Uniformity on the HypersphereTongzhou Wang, Phillip IsolaICML 2020 · 2,360 citations
- MIND: A Large-scale Dataset for News RecommendationFangzhao Wu, Ying Qiao, Jiun-Hung Chen, Chuhan Wu et al.ACL 2020 · 454 citations
- MixText: Linguistically-Informed Interpolation of Hidden Space for Semi-Supervised Text ClassificationJiaao Chen, Zichao Yang, Diyi YangACL 2020 · 340 citations
- COCO-LM: Correcting and Contrasting Text Sequences for Language Model PretrainingYu Meng, Chenyan Xiong, Payal Bajaj, Saurabh Tiwary et al.NeurIPS 2021 · 231 citations
Related papers
- Unsupervised Story Discovery from Continuous News Streams via Scalable Thematic EmbeddingSusik Yoon, Dongha Lee, Yunyi Zhang, Jiawei HanSIGIR 2023 · 8 citations
- HRSTORY: Historical News Review Based Online Story DiscoveryRenjie Zhou, Haoran Ye, Jian Wan, Yong LiaoKDD 2025
- Continual Learning on Noisy Data Streams via Self-Purified ReplayChris Dongjoo Kim, Jinseo Jeong, Sangwoo Moon, Gunhee KimICCV 2021 · 53 citations
- Self-Supervised Continual Graph Learning via Adaptive Spaced Replay on Node ProxiesZhen Peng, Xu Hua, Jingchen Hao, Qika Lin et al.KDD 2025
- NewsEmbed: Modeling News through Pre-trained Document RepresentationsJialu Liu, Tianqi Liu, Cong YuKDD 2021 · 15 citations
