HRSTORY: Historical News Review Based Online Story Discovery
Renjie Zhou, Haoran Ye, Jian Wan, Yong Liao
Abstract
Story discovery on news streams can help people quickly find story from vast amounts of news, improving the efficiency of information acquisition. Recent online story discovery methods encode text topics and then cluster articles into stories based on similarity. However, the results obtained by these methods are one-time, and clustered news cannot adaptively update in a continuous news stream. Additionally, the inadequate quality of article encoding and the presence of noise data deteriorate the performance of story discovery. To this end, we propose HRSTORY for online story discovery on news streams, which employs a historical news review method to enable news to continuously adapt to the latest environment in the stream data and make corrections and updates. Furthermore, HRSTORY captures better article embeddings through modeling multi-layer relational dependencies within the text. By using sentence-level noise masking, HRSTORY improves the relevance of news article representation to core topics and reduces the interference of noise data. Experiments on real news datasets show that HRSTORY outperforms the state-of-the-art algorithms in unsupervised online story discovery performance.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get d58caffd-4d31-4500-96d8-4cb8a16e0325Related papers
- SCStory: Self-supervised and Continual Online Story DiscoverySusik Yoon, Yu Meng, Dongha Lee, Jiawei HanWWW 2023 · 14 citations
- Unsupervised Story Discovery from Continuous News Streams via Scalable Thematic EmbeddingSusik Yoon, Dongha Lee, Yunyi Zhang, Jiawei HanSIGIR 2023 · 8 citations
- PDSum: Prototype-driven Continuous Summarization of Evolving Multi-document Sets StreamSusik Yoon, Hou Pong Chan, Jiawei HanWWW 2023 · 13 citations
- Be Relevant, Non-Redundant, and Timely: Deep Reinforcement Learning for Real-Time Event SummarizationMin Yang, Chengming Li, Fei Sun, Zhou Zhao et al.AAAI 2020 · 8 citations
- Deep Exogenous and Endogenous Influence Combination for Social Chatter Intensity PredictionSubhabrata Dutta, Sarah Masud, Soumen Chakrabarti, Tanmoy ChakrabortyKDD 2020 · 2 citations
