LOKI: Proactively Discovering Online Scam Websites by Mining Toxic Search Queries
Pujan Paudel, Gianluca Stringhini
Abstract
Online e-commerce scams, ranging from shopping scams to pet scams, globally cause millions of dollars in financial damage every year. In response, the security community has developed highly accurate detection systems able to determine if a website is fraudulent. However, finding candidate scam websites that can be passed as input to these downstream detection systems is challenging: relying on user reports is inherently reactive and slow, and proactive systems issuing search engine queries to return candidate websites suffer from low coverage and do not generalize to new scam types. In this paper, we present LOKI, a system designed to identify search engine queries likely to return a high fraction of fraudulent websites. LOKI implements a keyword scoring model grounded in Learning Under Privileged Information (LUPI) and feature distillation from Search Engine Result Pages (SERPs). We rigorously validate LOKI across 10 major scam categories and demonstrate a 20.58 times improvement in discovery over both heuristic and data-driven baselines across all categories. Leveraging a small seed set of only 1,663 known scam sites, we use the keywords identified by our method to discover 52,493 previously unreported scams in the wild. Finally, we show that LOKI generalizes to previously-unseen scam categories, highlighting its utility in surfacing emerging threats.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 056f2e1f-63e2-4034-8c6a-c7ed53970f6eBuilds on9
- Tranco: A Research-Oriented Top Sites Ranking Hardened Against ManipulationVictor Le Pochat, Tom van Goethem, Samaneh Tajalizadehkhoob, Maciej Korczynski et al.NDSS 2019 · 826 citations
- One-for-All: Bridge the Gap Between Heterogeneous Architectures in Knowledge DistillationZhiwei Hao, Jianyuan Guo, Kai Han, Yehui Tang et al.NeurIPS 2023 · 205 citations
- Less Defined Knowledge and More True Alarms: Reference-based Phishing Detection without a Pre-defined Reference ListRuofan Liu, Yun Lin, Xiwen Teoh, Gongshen Liu et al.USENIX Security 2024 · 39 citations
- The Poorest Man in Babylon: A Longitudinal Study of Cryptocurrency Investment ScamsMuhammad Muzammil, Abisheka Pitumpe, Xigao Li, Amir Rahmati et al.WWW 2025 · 13 citations
- Enabling Contextual Soft Moderation on Social Media through Contrastive Textual DeviationPujan Paudel, Mohammad Hammas Saeed, Rebecca Auger, Chris Wells et al.USENIX Security 2024 · 3 citations
Related papers
- Beyond Phish: Toward Detecting Fraudulent e-Commerce Websites at ScaleMarzieh Bitaab, Haehyun Cho, Adam Oest, Zhuoer Lyu et al.S&P 2023
- NOKEScam: Understanding and Rectifying Non-Sense Keywords Spear Scam in Search EnginesMingxuan Liu, Yunyi Zhang, Lijie Wu, Baojun Liu et al.USENIX Security 2025
- Detecting Credential Spearphishing in Enterprise SettingsGrant Ho, Aashish Sharma, Mobin Javed, Vern Paxson et al.USENIX Security 2017 · 94 citations
- Measuring Real-World Prompt Injection Attacks in LLM-based Resume ScreeningMohan Zhang, Yuqi Jia, Zhen Tan, Steven Jiang et al.USENIX Security 2026 · 4 citations
- Surveylance: Automatically Detecting Online Survey ScamsAmin Kharraz, William K. Robertson, Engin KirdaS&P 2018 · 73 citations
