Specious Sites: Tracking the Spread and Sway of Spurious News Stories at Scale
Hans W. A. Hanley, Deepak Kumar, Zakir Durumeric
Abstract
Misinformation, propaganda, and outright lies proliferate on the web, with some narratives having dangerous real-world consequences on public health, elections, and individual safety. However, despite the impact of misinformation, the research community largely lacks automated and programmatic approaches for tracking news narratives across online platforms. In this work, utilizing daily scrapes of 1,334 unreliable news websites, the large-language model MPNet, and DP-Means clustering, we introduce a system to automatically identify and track the narratives spread within online ecosystems. Identifying 52,036 narratives on these 1,334 websites, we describe the most prevalent narratives spread in 2022 and identify the most influential websites that originate and amplify narratives. Finally, we show how our system can be utilized to detect new narratives originating from unreliable news websites and to aid fact-checkers in more quickly addressing misinformation. We release code and data at https://github.com/hanshanley/specious-sites.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- On the Characteristics and Impacts of Protestware LibrariesTanner Finken, Jesse Chen, Sazzadur RahamanFSE 2025
- Hierarchical Level-Wise News Article Clustering via Multilingual Matryoshka EmbeddingsHans William Alexander Hanley, Zakir DurumericACL 2025
- Twits, Toxic Tweets, and Tribal Tendencies: Trends in Politically Polarized Posts on TwitterHans W. A. Hanley, Zakir DurumericCSCW 2025
- Characterizing and Detecting Propaganda-Spreading Accounts on TelegramKlim Kireev, Yevhen Mykhno, Carmela Troncoso, Rebekah OverdorfUSENIX Security 2025
- Cross-Platform Narrative Prediction: Leveraging Platform-Invariant Discourse NetworksPatrick Gerard, Luca Luceri, Leonardo Blas, Emilio FerraraWWW 2026
Builds on14
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 2,496 citations
- Understanding the Mirai BotnetManos Antonakakis, Tim April, Michael D. Bailey, Matt Bernhard et al.USENIX Security 2017 · 2,003 citations
- MPNet: Masked and Permuted Pre-training for Language UnderstandingKaitao Song, Xu Tan, Tao Qin, Jianfeng Lu et al.NeurIPS 2020 · 1,957 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding SharingPengcheng He, Jianfeng Gao, Weizhu ChenICLR 2023 · 394 citations
Related papers
- Tracking the Takes and Trajectories of English-Language News Narratives across Trustworthy and Worrisome WebsitesHans W. A. Hanley, Emily Okabe, Zakir DurumericUSENIX Security 2025
- DiNaM: Disinformation Narrative Mining with Large Language ModelsWitold Sosnowski, Arkadiusz Modzelewski, Kinga Skorupska, Adam WierzbickiEMNLP 2025
- PolyNarrative: A Multilingual, Multilabel, Multi-domain Dataset for Narrative Extraction from News ArticlesNikolaos Nikolaidis, Nicolas Stefanovitch, Purificação Silvano, Dimitar Iliyanov Dimitrov et al.ACL 2025
- Mitigating Adversarial Attacks by Transferring LLM-generated Narrative Reasoning for Robust Fake News DetectionMengyang Chen, Lingwei Wei, Wei Zhou, Songlin HuSIGIR 2026
- DiNO: Disinformation Narrative ObserverWitold Sosnowski, Arkadiusz Modzelewski, Kinga Skorupska, Adam WierzbickiACL 2026
